<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">25073</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2022.025073</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Practical Machine Learning Techniques for COVID-19 Detection Using Chest X-Ray Images</article-title><alt-title alt-title-type="left-running-head">Practical Machine Learning Techniques for COVID-19 Detection Using Chest X-Ray Images</alt-title><alt-title alt-title-type="right-running-head">Practical Machine Learning Techniques for COVID-19 Detection Using Chest X-Ray Images</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Mangalmurti</surname><given-names>Yurananatul</given-names></name>
</contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wattanapongsakorn</surname><given-names>Naruemon</given-names></name><email>naruemon.wat@kmutt.ac.th</email>
</contrib>
<aff id="aff-1"><institution>Department of Computer Engineering, King Mongkut&#x2019;s University of Technology Thonburi</institution>, <addr-line>Bangkok, 10140</addr-line>, <country>Thailand</country></aff>
</contrib-group><author-notes><corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Naruemon Wattanapongsakorn. Email: <email>naruemon.wat@kmutt.ac.th</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-05-02"><day>02</day>
<month>05</month>
<year>2022</year></pub-date>
<volume>34</volume>
<issue>2</issue>
<fpage>733</fpage>
<lpage>752</lpage>
<history>
<date date-type="received"><day>10</day><month>11</month><year>2021</year></date>
<date date-type="accepted"><day>07</day><month>1</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Mangalmurti and Wattanapongsakorn</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Mangalmurti and Wattanapongsakorn</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_25073.pdf"></self-uri>
<abstract>
<p>This paper presents effective techniques for automatic detection/classification of COVID-19 and other lung diseases using machine learning, including deep learning with convolutional neural networks (CNN) and classical machine learning techniques. We had access to a large number of chest X-ray images to use as input data. The data contains various categories including COVID-19, Pneumonia, Pneumothorax, Atelectasis, and Normal (without disease). In addition, chest X-ray images with many findings (abnormalities and diseases) from the National Institutes of Health (NIH) was also considered. Our deep learning approach used a CNN architecture with VGG16 and VGG19 models which were pre-trained with ImageNet. We compared this approach with the classical machine learning approaches, namely Support Vector Machine (SVM) and Random Forest. In addition to independently extracting image features, pre-trained features obtained from a VGG19 model were utilized with these classical machine learning techniques. Both binary and categorical (multi-class) classification tasks were considered on classical machine learning and deep learning. Several X-ray images ranging from 7000 images up to 11500 images were used in each of our experiments. Five experimental cases were considered for each classification model. Results obtained from all techniques were evaluated with confusion matrices, accuracy, precision, recall and F1-score. In summary, most of the results are very impressive. Our deep learning approach produced up to 97.5&#x0025; accuracy and 98&#x0025; F1-score on COVID-19 <italic>vs.</italic> non-COVID-19 (normal or diseases excluding COVID-19) class, while in classical machine learning approaches, the SVM with pre-trained features produced 98.9&#x0025; accuracy, and at least 98.2&#x0025; precision, recall and F1-score on COVID-19 <italic>vs.</italic> non-COVID-19 class. These disease detection models can be deployed for practical usage in the near future.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>COVID-19</kwd>
<kwd>deep learning</kwd>
<kwd>image classification</kwd>
<kwd>lung disease</kwd>
<kwd>machine learning</kwd>
<kwd>pneumonia</kwd>
<kwd>pretrained features</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In late 2019, China experiences an outbreak of a new disease called SARS-CoV-2 or COVID-19. Among other symptoms, the disease can cause serious damage to the lungs of infected patients. The number of patients has been increasing rapidly. As of October 2021, over 240 million cases have been detected [<xref ref-type="bibr" rid="ref-1">1</xref>]. Of all the cases, there are 4.9 million death cases. The World Health Organization (WHO) has declared this disease a pandemic. So, there is an urgent need for an easy and simple way to detect the suspected COVID-19 symptoms as soon as possible. Chest X-ray image analysis is one of the main methods to detect lung disease including COVID-19. The image detection requires a radiologist together with a medical doctor to perform a diagnosis.</p>
<p>Artificial intelligence (AI) has shown great promise in image identification and classification tasks. Machine learning is one AI area that includes many methods for image detection and classification. One of the popular methods is called deep learning. It is built with a multi-layer network that can be trained using input data. A convolutional neural network (CNN) is one of the deep learning methods. It has great success in classifying spatial data is and usually used for recognizing images. CNN combines feature extraction and classification in one workflow. Other machine learning techniques require feature extraction as an additional step before performing detection/classification. However, the classical machine learning methods can also be successful in image classification.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p>Given the critical importance of COVID-19 diagnosis, there has been a wealth of recent research applying artificial intelligence for chest X-ray image diagnosis. Arias-Londo&#x00F1;o et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] used CNN for chest X-ray classification. The experimental corpus held many chest X-ray images in 3 classes which are COVID-19, pneumonia, and others as shown in <xref ref-type="table" rid="table-1">Tab. 1</xref>. Normal data (without disease) was not included. The best accuracy was obtained with 91.67&#x0025; accuracy. Khan et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed the classification using Xception CNN architecture with 71 layers deep. The dataset contained 290 COVID-19 images, 1203 normal images and 1653 pneumonia images. The best experiment used a binary classification model which obtained 99&#x0025; accuracy. Another CNN model was used with a decision tree classifier proposed by Vinod et al. [<xref ref-type="bibr" rid="ref-4">4</xref>]. The model produced 87&#x0025; of accuracy and 93&#x0025; of recall. Later, Abbas et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] proposed a framework combining Decomposition, Transfer learning, and Composition (DeTraC). The dataset contained a total of 196 X-ray images having 3 classes which were normal, COVID-19, and SARS. The final classification was done with VGG19 architecture. The model produced 97.35&#x0025; accuracy. Sedik et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] proposed CNN and other deep learning techniques. The methods were tested on a dataset of 56 COVID-19 and 56 non-COVID-19 images. The dataset was augmented by 10-fold before using in the model. This research produced up to 99&#x0025; accuracy. Multiple architectures of CNN were tested by Kamil [<xref ref-type="bibr" rid="ref-7">7</xref>]. All architectures used transfer learning techniques so that they could be learned quickly. The dataset combined 23 CT images and 977 Chest X-ray images (with 195 COVID-19 images) in 2 classes, normal and COVID-19. The VGG19 model gave the highest accuracy of 99&#x0025;. Rangarajan et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] proposed a classification model using 2 data augmentation methods. The first method used flipping and rotation of the image, while the second method used a generative adversarial network [<xref ref-type="bibr" rid="ref-9">9</xref>] to create synthetic images. The researchers compared 5 classification models, where the best model is VGG16 with accuracy of 98.6&#x0025;. Mor&#x00ED;s et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] improved the COVID-19 screening with portable chest X-ray images using cycle-generative adversarial networks to generate synthetic images. The classification task was to classify the COVID-19 class and non-COVID-19 class where the ResNet-9 model gave the highest accuracy of 98.61&#x0025;.</p>
<table-wrap id="table-1"><label>Table 1</label>
<caption>
<title>Summary of all literature review showing the methods, amount of data, number of output classes, and accuracy</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Reference</th>
<th align="left" rowspan="2">Methods used</th>
<th align="center" colspan="4">Cases</th>
<th align="left" rowspan="2">Output Classes</th>
<th align="left" rowspan="2">Accuracy (&#x0025;)</th>
</tr>
<tr>
<th align="left">Normal</th>
<th align="left">COVID</th>
<th align="left">Pneumonia</th>
<th align="left">Others</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td align="left">CNN</td>
<td align="left">-</td>
<td align="left">8573</td>
<td align="left">24114</td>
<td align="left">49983</td>
<td align="left">3</td>
<td align="left">up to 91.6</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td align="left">Xception CNN</td>
<td align="left">1203</td>
<td align="left">290</td>
<td align="left">1653</td>
<td align="left">-</td>
<td align="left">2&#x2013;4</td>
<td align="left">up to 99</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td align="left">CNN &#x002B; DT</td>
<td align="left">463</td>
<td align="left">701</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 87</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-5">5</xref>]</td>
<td align="left">VGG19</td>
<td align="left">80</td>
<td align="left">105</td>
<td align="left">-</td>
<td align="left">11</td>
<td align="left">3</td>
<td align="left">up to 97.3</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td align="left">CNN</td>
<td align="left">56</td>
<td align="left">56</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 99</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td align="left">VGG19</td>
<td align="left">805</td>
<td align="left">195</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 99</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td align="left">VGG16</td>
<td align="left">1304</td>
<td align="left">598</td>
<td align="left">3804</td>
<td align="left">-</td>
<td align="left">3</td>
<td align="left">up to 98.6</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td align="left">ResNet-9</td>
<td align="left">240</td>
<td align="left">240</td>
<td align="left">240</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 98.6</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td align="left">Majority Vote</td>
<td align="left">782</td>
<td align="left">782</td>
<td align="left">782</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 98</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td align="left">SVM</td>
<td align="left">234</td>
<td align="left">87</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 100</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td align="left">SVM</td>
<td align="left">24</td>
<td align="left">101</td>
<td align="left">24</td>
<td align="left">111</td>
<td align="left">6</td>
<td align="left">up to 94.2</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td align="left">Ensemble</td>
<td align="left">-</td>
<td align="left">306</td>
<td align="left">306</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 93.8</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td align="left">XGBoost</td>
<td align="left">1341</td>
<td align="left">206</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">2</td>
<td align="left">up to 98.7</td>
</tr>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td align="left">SVM</td>
<td align="left">75</td>
<td align="left">341</td>
<td align="left">75</td>
<td align="left">52</td>
<td align="left">2</td>
<td align="left">up to 95.2</td>
</tr>
<tr>
<td align="left">Ours</td>
<td align="left">SVM, RF, VGG16, VGG19</td>
<td align="left">3500</td>
<td align="left">3500</td>
<td align="left">1500</td>
<td align="left">3000</td>
<td align="left">2&#x2013;5</td>
<td align="left">up to 99</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Many recent research publications have presented classical machine learning techniques for COVID-19 chest X-ray image detection/classification as well. Majority voting-based techniques were deployed to predict the result by Chandra et al. [<xref ref-type="bibr" rid="ref-11">11</xref>], where classical machine learning models were used to predict the disease classes of chest X-ray images. This research used 3 feature extraction methods to select the significant features [<xref ref-type="bibr" rid="ref-12">12</xref>]. The research obtained 98&#x0025; accuracy on normal <italic>vs.</italic> abnormal, and 91&#x0025; accuracy on pneumonia <italic>vs.</italic> COVID-19. Tuncer et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] used Local Binary Pattern feature extraction [<xref ref-type="bibr" rid="ref-14">14</xref>] and ReliefF feature selection [<xref ref-type="bibr" rid="ref-15">15</xref>]. A Support Vector Machine (SVM) classifier obtained 100&#x0025; accuracy. &#x00D6;zt&#x00FC;rk [<xref ref-type="bibr" rid="ref-16">16</xref>] also presented an SVM classifier considering both CT X-ray and chest X-ray images using shrunken features. The best accuracy obtained was 94.23&#x0025;. Lastly, CT X-ray image classification was performed by Ardakani et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] using different techniques, where the dataset used only COVID-19 and pneumonia images with a total of 612 images. The research presented 5 classification models, where the best model was an ensemble model using 4 of the models to predict the result with accuracy of 93.85&#x0025;. J&#x00FA;nior et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] used deep features from 3 architectures, VGG19, Inception-v3, and ResNet50 as input of XGBoost classifier. The performance of classifying normal and COVID-19 classes gave 98.71&#x0025; accuracy. Tamal et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] used radiomic features as input in SVM classifier. The radiomic features are special features extracted from radiographic medical images [<xref ref-type="bibr" rid="ref-20">20</xref>]. The classification output classes are COVID-19 and non-COVID-19 with 95.2&#x0025; accuracy. <xref ref-type="table" rid="table-1">Tab. 1</xref> summarizes these previous research works.</p>
<p>In summary, both classical machine learning techniques and deep learning techniques were considered for COVID-19 detection/classification. However, most of this previous research used a small number of COVID-19 chest X-ray images together with a few hundreds of normal images. Mainly, they considered binary class classification. A few papers considered 3-class classification using input data containing a few disease types; COVID-19 and pneumonia.</p>
<p>In our work, we want to be able to identify multiple lung diseases, particularly COVID-19 against others. We use a large amount of data (at least 1500 images per disease) compared to the previous research to improve the detection/classification results and make them more reliable and robust. Multiple reliable datasets with various lung diseases are applied. We also want to directly compare the detection/classification performance of deep learning and classical machine learning techniques. Since some research papers have reported only overall detection accuracy, we want to perform a more extensive evaluation that considers the full confusion matrix as well as other measures. Finally, we want to compare results from binary class classification with multi-class classification. The framework of this work is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed method framework</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-1.png"/>
</fig>
<p>The main contributions of this paper are as follows.<list list-type="simple"><list-item><label>&#x25A0;</label>
<p>Consider and evaluate both deep learning techniques and classical machine learning (ML) techniques for binary-class and multi-class classification. Many experimental cases are presented.</p></list-item><list-item><label>&#x25A0;</label>
<p>Apply pre-trained CNN models to speed up the training process and produce high detection/classification accuracy.</p></list-item><list-item><label>&#x25A0;</label>
<p>Apply pre-trained features from deep learning models to classical ML techniques to produce higher detection/classification accuracy than when using previously proposed feature extraction and selection techniques.</p></list-item><list-item><label>&#x25A0;</label>
<p>Provide up to 97.5&#x0025; accuracy and 98&#x0025; F1-score of COVID-19 detection against various lung diseases and normal using the proposed deep learning models.</p></list-item><list-item><label>&#x25A0;</label>
<p>Provide up to 98.9&#x0025; accuracy and 99&#x0025; F1-score of COVID-19 detection using the proposed Support Vector Machine with pre-trained features from a selected deep learning model.</p></list-item></list></p>
<p>This paper is organized as follows. Section 2 presents the proposed methods including dataset preview, data preprocessing, and detection/classification techniques. Section 3 describes our experimental design, parameter settings, evaluation methods, and performance parameters. Section 4 presents the obtained results and Section 5 provides the conclusions.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Method</title>
<p>In this section, we describe our datasets, how we process them, and all methods used in experiments. We also discuss the deep learning and classical machine learning (ML) techniques used in this paper.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset</title>
<p>Our dataset consists of chest X-ray images with Posterior Anterior (PA) view as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> (more detailed information is provided in Section 3.). We obtained the datasets from reliable online sources such as Kaggle and NIH databases [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. The datasets contain 5 classes which are Atelectasis, COVID-19, Normal, Pneumonia, and Pneumothorax.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Sample images from the dataset containing atelectasis, COVID-19, normal, pneumonia, and pneumothorax classes</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-2.png"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Preprocessing</title>
<p>Chest X-ray images were first in grayscale format and then converted into red-green-blue (RGB) format for processing with a deep learning model and feature extraction method. The images were 224 &#x00D7; 224 with 3 color channels in RGB format. For each class, we mapped the images to their labels. For binary-class classifications, 0 and 1 are used as our output. For categorical or multi-class classification, one-hot encoding is used to transform the numeric label into values that represent each class. Normal means having no lung disease.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Deep Learning</title>
<p>Deep learning is one of machine learning techniques that uses artificial neural network as its core. Convolutional Neural Network (CNN) is a well-known technique that is used to classify images. Deep learning usually consists of many layers. The network uses a back-propagation learning algorithm to adapt the weights linking neurons in the network to one another. The weights determine the final output of the network when given an input to classify. Deep learning already contains feature extraction, unlike classical machine learning that requires an additional feature extraction process before classification as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Comparison between classical machine learning and deep learning process</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-3.png"/>
</fig>
<p>Convolutional neural networks (CNN) [<xref ref-type="bibr" rid="ref-24">24</xref>] are specialized for handling spatially distributed data such as images. They contain convolutional layers in which the output from a particular neuron or unit depends not only on its activation value but also on its position in an array, with respect to other units. The convolutional layers of a CNN are primarily responsible for learning spatial patterns, that is, features. To work well, CNNs also include other types of layers, including pooling layers, drop-out layers, and fully connected layers.</p>
<p>The convolutional layer is used to collect information from image inputs and recombine them into new images or features that will be fed into the deeper layer. The convolution layers are the core of the convolutional neural network. The convolution operation is mathematically defined in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>, where <italic>G</italic> is feature mapping, <italic>h</italic> is a kernel, and <italic>f</italic> is an input image. Pooling layers are used to reduce the dimensionality of the network and summarize the output from convolutional layers. Dropout layers randomly discard some of the output weights from the layer. Dropout layers slow the rate of learning but help make the network generalize better by reducing overfitting. The fully interconnected layer (FC) is the last part of the network. It is used to classify the input and gives predicted output. The input will be converted from two-dimensional arrays to a one-dimensional vector (called &#x201C;flattened&#x201D;) before entering this layer.<disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi>G</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mi>j</mml:mi></mml:munder><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:munder><mml:mo>&#x2061;</mml:mo><mml:mi>h</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">j</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">k</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>f</mml:mi><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">m</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">j</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">k</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math>
</disp-formula></p>
<p>A popular method used in training a deep learning model is called &#x201C;transfer learning&#x201D; [<xref ref-type="bibr" rid="ref-25">25</xref>]. The method adapts the knowledge of an already trained model for a new training goal. For classification, the transfer learning freezes the trained capability of a previously trained model and removes the classification layers which are fully connected layers. Then, we replace them with new fully connected layers and train the model further to reach our goal. The transfer learning uses the lower-level features learned by the original model, allowing them to be refined and applied in a new model to solve a new problem. Many successful architectures have used the ImageNet training dataset [<xref ref-type="bibr" rid="ref-26">26</xref>] and gained high detection performance. VGG is one of them. It has 2 main variation models which are VGG16 and VGG19 [<xref ref-type="bibr" rid="ref-27">27</xref>]. The VGG16 model has 16 layers while the VGG19 model has 19 layers as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The layer counting does not include max pooling layers and fully connected layers.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Original architecture of VGG16 and VGG19</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-4.png"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Feature Extraction</title>
<p>Unlike CNN, classical machine learning techniques require the researcher to define and extract features (dimensions or attributes) from the input before training the model. Feature extraction calculates values from the input that represent the important characteristics of the data. In this work, we use pre-trained features to extract the feature from our data.</p>
<p>The pre-trained feature approach borrows the idea from the transfer learning model with trained weights. Usually, transfer learning is applied to a new CNN model, but in our case, the deep-learning-trained features are used with classical ML approaches. We start with the VGG19 model trained with the ImageNet dataset. The fully connected layers in the model are removed, so that the last layer is the max-pooling layer, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The model receives input size 224 &#x00D7; 224 &#x00D7; 3 and gives an output size of 7 &#x00D7; 7 &#x00D7; 512. Them, the output features are flattened to produce 25,088 features and will be later used as the input to the classical machine learning techniques.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Cut-off VGG19 architecture to output only features</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-5.png"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Classical Machine Learning</title>
<p>Like deep learning models, supervised classical machine learning techniques must be trained with labeled samples of the input data. However, unlike deep learning, these techniques do not learn the best features in the input dataset by themselves. Instead, the input data must be transformed into a set of selected features, for each training and testing data item. In this work, we present two effective machine learning techniques which are Support Vector Machine (SVM) and Random Forest (RF).</p>
<p>SVM [<xref ref-type="bibr" rid="ref-28">28</xref>] is a machine learning technique that separates data items (input data items, represented as vectors of feature values) into classes in high dimensional space by using a hyperplane. The data items in the space can be handled with a selected kernel function that maps the data into a different form. There are 3 common kernel functions which are linear, polynomial, and Gaussian functions.</p>
<p>RF [<xref ref-type="bibr" rid="ref-29">29</xref>] is an ensemble technique that consists of multiple decision trees [<xref ref-type="bibr" rid="ref-30">30</xref>]. The RF method uses a randomly selected subset of data features to create a tree and creates many trees instead of one. The data is classified based on the result of many trees, using some aggregation algorithm.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Design</title>
<p>In this section, the input data used in each experiment is described together with the experimental setup, evaluation method, performance evaluation metrics, and hardware and software environment for the experiments.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset</title>
<p>The chest X-ray dataset used in our experiments consists of 5 classes which are normal, COVID-19, pneumonia, pneumothorax, and atelectasis. These diseases are well-known and commonly found all around the world. The input images were obtained from online databases at Kaggle (kaggle.com) and the National Institutes of Health (nihcc.app.box.com). Details of the datasets are shown in <xref ref-type="table" rid="table-2">Tab. 2</xref>.</p>
<table-wrap id="table-2"><label>Table 2</label>
<caption>
<title>Number of images gathered for each class</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Class</th>
<th align="left">Kaggle</th>
<th align="left">NIH</th>
<th align="left">Total</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Normal</td>
<td align="left">-</td>
<td align="left">3,500</td>
<td align="left">3,500</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left">3,500</td>
<td align="left">-</td>
<td align="left">3,500</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left">1,500</td>
<td align="left">-</td>
<td align="left">1,500</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">-</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">-</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Parameter Settings</title>
<p>The parameter setting for deep learning and classical machine learning are presented in <xref ref-type="table" rid="table-3">Tabs. 3</xref> and <xref ref-type="table" rid="table-4">4</xref>, respectively. The deep learning models that we consider are VGG 16 and VGG19, while the machine learning models are SVM and RF.</p>
<table-wrap id="table-3"><label>Table 3</label>
<caption>
<title>Parameter setting for deep learning models, VGG16 and VGG19</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Parameter</th>
<th align="left">VGG16</th>
<th align="left">VGG19</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Base model</td>
<td align="left">VGG16 without fully connected layers</td>
<td align="left">VGG19 without fully connected layers</td>
</tr>
<tr>
<td align="left">Input shape</td>
<td align="left">224&#x2009;&#x00D7;&#x2009;224&#x2009;&#x00D7;&#x2009;3</td>
<td align="left">224&#x2009;&#x00D7;&#x2009;224&#x2009;&#x00D7;&#x2009;3</td>
</tr>
<tr>
<td align="left">FC 1</td>
<td align="left">1024</td>
<td align="left">1024</td>
</tr>
<tr>
<td align="left">FC 2</td>
<td align="left">1024</td>
<td align="left">1024</td>
</tr>
<tr>
<td align="left">FC 3</td>
<td align="left">512</td>
<td align="left">512</td>
</tr>
<tr>
<td align="left">Output for binary class</td>
<td align="left">Sigmoid</td>
<td align="left">Sigmoid</td>
</tr>
<tr>
<td align="left">Output for categorical classes</td>
<td align="left">SoftMax</td>
<td align="left">SoftMax</td>
</tr>
<tr>
<td align="left">Optimizer</td>
<td align="left">ADAM</td>
<td align="left">ADAM</td>
</tr>
<tr>
<td align="left">Batch size</td>
<td align="left">32</td>
<td align="left">32</td>
</tr>
<tr>
<td align="left">Epochs</td>
<td align="left">100</td>
<td align="left">100</td>
</tr>
<tr>
<td align="left">Learning rate</td>
<td align="left">1E-7</td>
<td align="left">1E-7</td>
</tr>
<tr>
<td align="left">Initial weight</td>
<td align="left">ImageNet</td>
<td align="left">ImageNet</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-4"><label>Table 4</label>
<caption>
<title>Parameter setting for SVM and RF</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Algorithm</th>
<th align="left">Parameter</th>
<th align="left">Input size</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">SVM</td>
<td align="left">kernel &#x003D; Gaussian</td>
<td align="left" rowspan="2">25088<break/>25088</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="left">depth&#x2009;&#x003D;&#x2009;16, estimators&#x2009;&#x003D;&#x2009;250</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental Design</title>
<p>We considered both binary (2 classes) and categorical (3&#x2013;5 classes) classification tasks. For the binary classification, two pairwise discriminations: normal <italic>vs.</italic> COVID-19 (case 1) and COVID-19 <italic>vs.</italic> non-COVID-19 (case 2) were examined. For the categorical classification, 3 classes (normal, COVID-19 and others) and 5 classes (normal, COVID-19, pneumonia, pneumothorax, and atelectasis) were considered. <xref ref-type="table" rid="table-5">Tab. 5</xref> shows the detailed experimental design and input dataset for each experimental case.</p>
<table-wrap id="table-5"><label>Table 5</label>
<caption>
<title>Input data on each experimental case</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Case</th>
<th align="left" rowspan="2">Experiment</th>
<th align="left" colspan="5">Classes</th>
</tr>
<tr>
<th align="left">Normal</th>
<th align="left">COVID</th>
<th align="left">Pneumonia</th>
<th align="left">Pneumothorax</th>
<th align="left">Atelectasis</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">COVID-19 <italic>vs.</italic> normal</td>
<td align="left">3,500</td>
<td align="left">3,500</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">-</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">COVID-19 <italic>vs.</italic> non-COVID-19</td>
<td align="left">1,500</td>
<td align="left">3,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">COVID-19 <italic>vs.</italic> normal <italic>vs.</italic> others</td>
<td align="left">3,500</td>
<td align="left">3,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">All 5 classes</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
<td align="left">1,500</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Evaluation</title>
<p>For evaluating our detection/classification models, we use two methods to separate the training and testing data which are a hold-out method and a k-fold cross-validation method. In the deep learning experiments, we used only the hold-out method. The method splits the training, testing, and validation data into three non-overlapping sets. Validation data is only used to observe the result from model training to find the best time to stop training to avoid over-trained or overfitting. K-fold cross-validation splits the data into k subsets. Then it tests each subset while using the model trained from the remaining subsets. In our experiment, the k-fold method is only used in classical machine learning. The detailed information of our evaluation is shown in <xref ref-type="table" rid="table-6">Tab. 6</xref>.</p>
<table-wrap id="table-6"><label>Table 6</label>
<caption>
<title>Evaluation for each classifier</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Classifier</th>
<th align="left">Train</th>
<th align="left">Test</th>
<th align="left">Validation</th>
<th align="left">k-fold</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">VGG16</td>
<td align="left">70&#x0025;</td>
<td align="left">20&#x0025;</td>
<td align="left">10&#x0025;</td>
<td align="left">-</td>
</tr>
<tr>
<td align="left">VGG19</td>
<td align="left">70&#x0025;</td>
<td align="left">20&#x0025;</td>
<td align="left">10&#x0025;</td>
<td align="left">-</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">5</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">5</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Performance Metrics</title>
<p>Performance evaluation metrics are calculated from the confusion matrix including accuracy, precision, recall, and F1-score. The definitions of the confusion matrix and performance evaluation metrics are shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> and <xref ref-type="table" rid="table-7">Tab. 7</xref>, respectively, where TP is true positive, FP is false positive, FN is false negative, and TN is True negative.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Confusion matrix of 2 classes</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-6.png"/>
</fig>
<table-wrap id="table-7"><label>Table 7</label>
<caption>
<title>Performance evaluation for each classifier</title></caption>
<table><colgroup><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Measurement</th>
<th align="left">Formula</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Recall</td>
<td align="left">TP/(TP &#x002B; FN)</td>
</tr>
<tr>
<td align="left">Precision</td>
<td align="left">TP/(TP &#x002B;FP)</td>
</tr>
<tr>
<td align="left">Accuracy</td>
<td align="left">(TP &#x002B; TN)/(TP &#x002B; TN &#x002B; FP &#x002B; FN)</td>
</tr>
<tr>
<td align="left">F1-score</td>
<td align="left">(2 &#x002A; Precision &#x002A; Recall)/(Precision &#x002B; Recall)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Environment</title>
<p>In this paper, all detection/classification models were developed/trained using a personal computer with i5-9500F 3.00&#x2005;GHz Processor, 16 GB RAM, and GTX 1060 Graphics card running on the operating system Windows 10, 64-bit Pro. We conducted the experiments using Python 3 with Keras deep learning library and Scikit learn classical machine learning library.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Results</title>
<p>In this section, all experimental results of our proposed methods are presented. The results were separated into two parts: deep learning experiments, and classical ML experiments with pre-trained features. Each part contained 4 experimental cases (binary and categorical) as previously discussed.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Deep Learning Experiment</title>
<p>Our first part of the experiments focused on pure deep learning methods using transfer learning with ImageNet weight. The experiment contained 4 cases which are Normal <italic>vs.</italic> COVID-19, COVID-19 <italic>vs.</italic> non-COVID-19, Normal <italic>vs.</italic> COVID-19 <italic>vs.</italic> Others, and all 5 classes. As discussed in Section 3.4, the experimental cases split the train-test-validation in the ratio of 70:20:10 using the hold-out method. The training/validation accuracy and training/validation loss results of VGG19 are graphically displayed in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Accuracy and loss values of VGG19 from cases 1 to 4 starting from top to bottom rows</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-7.png"/>
</fig>
<p>From <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, we see that the accuracy of both training and validation for all models increase rapidly during the first 10 iterations/epochs and then gradually increase. The first row presents result obtained from the first experimental case, while the 2<sup>nd</sup>-4<sup>th</sup> rows present results obtained from the 2<sup>nd</sup>-4<sup>th</sup> cases, respectively. During this training time, each classification model learns to adjust itself so that the difference between the two detection/accuracy rates (training and validation) is small and becomes stable. Then, the models are ready for testing with separate testing datasets, where results are presented with confusion matrices as shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, and the corresponding performance evaluation metrics are shown in <xref ref-type="table" rid="table-8">Tabs. 8</xref> and <xref ref-type="table" rid="table-9">9</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Results of 4 experimental cases: confusion matrices of VGG16 and VGG19 classification models started from the top row</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_25073-fig-8.png"/>
</fig>
<table-wrap id="table-8"><label>Table 8</label>
<caption>
<title>Performance of VGG16</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Experiments</th>
<th align="left" rowspan="2">Classes</th>
<th align="center" colspan="4">Performance Metrics</th>
</tr>
<tr>
<th align="left"><italic>Accuracy</italic></th>
<th align="left"><italic>Recall</italic></th>
<th align="left"><italic>Precision</italic></th>
<th align="left"><italic>F1-score</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="2">Case 1: COVID-19 <italic>vs.</italic> normal</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2"><bold>0.965</bold></td>
<td align="left">0.975</td>
<td align="left">0.956</td>
<td align="left"><bold>0.966</bold></td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.955</td>
<td align="left">0.975</td>
<td align="left">0.965</td>
</tr>
<tr>
<td align="left" rowspan="2">Case 2: COVID-19 <italic>vs.</italic> non-COVID-19</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2"><bold>0.975</bold></td>
<td align="left">0.985</td>
<td align="left">0.976</td>
<td align="left"><bold>0.980</bold></td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.958</td>
<td align="left">0.973</td>
<td align="left">0.966</td>
</tr>
<tr>
<td align="left" rowspan="3">Case 3: COVID-19 <italic>vs.</italic> normal <italic>vs.</italic> others</td>
<td align="left">Normal</td>
<td align="left">0.663</td>
<td align="left">0.457</td>
<td align="left">0.448</td>
<td align="left">0.452</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.921</bold></td>
<td align="left">0.862</td>
<td align="left">0.877</td>
<td align="left">0.870</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.662</td>
<td align="left">0.567</td>
<td align="left">0.569</td>
<td align="left">0.568</td>
</tr>
<tr>
<td align="left" rowspan="5">Case 4: 5-class classification</td>
<td align="left">Normal</td>
<td align="left">0.846</td>
<td align="left">0.603</td>
<td align="left">0.619</td>
<td align="left">0.611</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.903</bold></td>
<td align="left">0.750</td>
<td align="left">0.762</td>
<td align="left">0.756</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.959</bold></td>
<td align="left">0.890</td>
<td align="left">0.905</td>
<td align="left">0.897</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.782</td>
<td align="left">0.470</td>
<td align="left">0.457</td>
<td align="left">0.463</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.786</td>
<td align="left">0.783</td>
<td align="left">0.467</td>
<td align="left">0.475</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-9"><label>Table 9</label>
<caption>
<title>Performance of VGG19</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Experiments</th>
<th align="left" rowspan="2">Classes</th>
<th align="center" colspan="4">Performance metrics</th>
</tr>
<tr>
<th align="left"><italic>Accuracy</italic></th>
<th align="left"><italic>Recall</italic></th>
<th align="left"><italic>Precision</italic></th>
<th align="left"><italic>F1-score</italic></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="2">Case 1: COVID-19 <italic>vs.</italic> normal</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2"><bold>0.965</bold></td>
<td align="left">0.970</td>
<td align="left">0.961</td>
<td align="left"><bold>0.965</bold></td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.961</td>
<td align="left">0.969</td>
<td align="left">0.965</td>
</tr>
<tr>
<td align="left" rowspan="2">Case 2: COVID-19 <italic>vs.</italic> non-COVID-19</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2"><bold>0.974</bold></td>
<td align="left">0.962</td>
<td align="left">0.967</td>
<td align="left"><bold>0.964</bold></td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.980</td>
<td align="left">0.978</td>
<td align="left">0.979</td>
</tr>
<tr>
<td align="left" rowspan="3">Case 3: COVID-19 <italic>vs.</italic> normal <italic>vs.</italic> others</td>
<td align="left">Normal</td>
<td align="left">0.689</td>
<td align="left">0.480</td>
<td align="left">0.489</td>
<td align="left">0.484</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.927</bold></td>
<td align="left">0.871</td>
<td align="left">0.889</td>
<td align="left"><bold>0.880</bold></td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.690</td>
<td align="left">0.620</td>
<td align="left">0.601</td>
<td align="left">0.610</td>
</tr>
<tr>
<td align="left" rowspan="5">Case 4: 5-class classification</td>
<td align="left">Normal</td>
<td align="left">0.858</td>
<td align="left">0.630</td>
<td align="left">0.649</td>
<td align="left">0.639</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.934</bold></td>
<td align="left">0.790</td>
<td align="left">0.871</td>
<td align="left"><bold>0.828</bold></td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.968</bold></td>
<td align="left">0.933</td>
<td align="left">0.909</td>
<td align="left"><bold>0.921</bold></td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.782</td>
<td align="left">0.466</td>
<td align="left">0.457</td>
<td align="left">0.462</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.798</td>
<td align="left">0.533</td>
<td align="left">0.495</td>
<td align="left">0.513</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, we can see that the classification results obtained from the VGG16 and VGG19 in each experimental case are similar. COVID-19 images can be mostly detected/classified. In particular, in cases 1&#x2013;2 with binary-class classification, less than 50 X-ray images out of 1400 images (700 COVID-19 images and 700 non-COVID-19 images) were misclassified. Normal images, shown in case 1, can be identified mostly as well. Considering multi-class classification, cases 3&#x2013;4, more misclassified images were obtained. However, we were still able to identify COVID-19 correctly mostly, apart from normal and other diseases, as shown in case 3. Then, with more output classes to classify, in case 4, we had most of the misclassified images belong to Pneumothorax and Atelectasis classes. The performance of the VGG16 and VGG19 can be considered in detail as presented in the following tables.</p>
<p>From <xref ref-type="table" rid="table-8">Tabs. 8</xref> and <xref ref-type="table" rid="table-9">9</xref>, experimental cases 1 and 2 which are binary class classification gave the most promising results with over 0.95 on the accuracy, recall, precision, and F1-score. For experimental cases 3 and 4, we consider more lung diseases for classification with 3 and 5 classes, respectively. Both of the experiments give good results on COVID-19 detection/classification. In experimental case 4 with 5-class classification, the performance in COVID-19 classification is 0.934 accuracy, while pneumonia classification is the best, with 0.968 accuracy.</p>
<p>In summary, both VGG16 and VGG19 models give a similar performance at binary-class classification. However, VGG19 performs slightly better than VGG16 at multi-class classification.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Classical Machine Learning Experiments</title>
<p>Next, we considered classical ML techniques with pre-trained features of VGG19 as input data. All experimental cases use 5-fold cross-validation method. The input images are in red-green-blue format. The performance evaluation metrics for SVM and RF are presented in <xref ref-type="table" rid="table-10">Tabs. 10</xref> and <xref ref-type="table" rid="table-11">11</xref>, respectively, where SVM was found to have superior performance among all classification techniques that we considered in all experiments.</p>
<table-wrap id="table-10"><label>Table 10</label>
<caption>
<title>Performance of SVM using VGG19 pre-trained features</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Experiments</th>
<th align="left" rowspan="2">Classes</th>
<th align="left" rowspan="2">Stats</th>
<th align="center" colspan="4">Performance metrics</th>
</tr>
<tr>
<th align="left">Accuracy</th>
<th align="left">Recall</th>
<th align="left">Precision</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="4">Case 1: COVID-19 <italic>vs.</italic> normal</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Max</td>
<td align="left" rowspan="2"><bold>0.990</bold></td>
<td align="left">0.997</td>
<td align="left">0.995</td>
<td align="left">0.990</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.995</td>
<td align="left">0.997</td>
<td align="left">0.990</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Avg</td>
<td align="left" rowspan="2">0.985</td>
<td align="left">0.988</td>
<td align="left">0.981</td>
<td align="left">0.985</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.981</td>
<td align="left">0.988</td>
<td align="left">0.985</td>
</tr>
<tr>
<td align="left" rowspan="4">Case 2: COVID-19 <italic>vs.</italic> non-COVID-19</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Max</td>
<td align="left" rowspan="2"><bold>0.991</bold></td>
<td align="left">0.988</td>
<td align="left">0.994</td>
<td align="left">0.987</td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.996</td>
<td align="left">0.993</td>
<td align="left">0.992</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Avg</td>
<td align="left" rowspan="2">0.987</td>
<td align="left">0.979</td>
<td align="left">0.985</td>
<td align="left">0.982</td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.991</td>
<td align="left">0.988</td>
<td align="left">0.989</td>
</tr>
<tr>
<td align="left" rowspan="6">Case 3: COVID-19 <italic>vs.</italic> normal <italic>vs.</italic> others</td>
<td align="left">Normal</td>
<td align="left" rowspan="3">Max</td>
<td align="left">0.778</td>
<td align="left">0.817</td>
<td align="left">0.608</td>
<td align="left">0.687</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.990</bold></td>
<td align="left">0.991</td>
<td align="left">0.986</td>
<td align="left">0.985</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.779</td>
<td align="left">0.636</td>
<td align="left">0.791</td>
<td align="left">0.690</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left" rowspan="3">Avg</td>
<td align="left">0.774</td>
<td align="left">0.764</td>
<td align="left">0.602</td>
<td align="left">0.673</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left">0.986</td>
<td align="left">0.981</td>
<td align="left">0.973</td>
<td align="left">0.977</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.776</td>
<td align="left">0.606</td>
<td align="left">0.773</td>
<td align="left">0.679</td>
</tr>
<tr>
<td align="left" rowspan="10">Case 4: 5-class classification</td>
<td align="left">Normal</td>
<td align="left" rowspan="5">Max</td>
<td align="left">0.929</td>
<td align="left">0.790</td>
<td align="left">0.884</td>
<td align="left">0.817</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.990</bold></td>
<td align="left">0.973</td>
<td align="left">0.979</td>
<td align="left">0.976</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.992</bold></td>
<td align="left">0.983</td>
<td align="left">0.986</td>
<td align="left">0.981</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.842</td>
<td align="left">0.636</td>
<td align="left">0.608</td>
<td align="left">0.615</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.861</td>
<td align="left">0.760</td>
<td align="left">0.637</td>
<td align="left">0.673</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left" rowspan="5">Avg</td>
<td align="left">0.896</td>
<td align="left">0.634</td>
<td align="left">0.800</td>
<td align="left">0.705</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.985</bold></td>
<td align="left">0.968</td>
<td align="left">0.959</td>
<td align="left">0.964</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.989</bold></td>
<td align="left">0.974</td>
<td align="left">0.973</td>
<td align="left">0.974</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.825</td>
<td align="left">0.597</td>
<td align="left">0.560</td>
<td align="left">0.577</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.849</td>
<td align="left">0.688</td>
<td align="left">0.608</td>
<td align="left">0.645</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-11"><label>Table 11</label>
<caption>
<title>Performance of RF using VGG19 pre-trained features</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Experiments</th>
<th align="left" rowspan="2">Class</th>
<th align="left" rowspan="2">Stats</th>
<th align="center" colspan="4">Performance metrics</th>
</tr>
<tr>
<th align="left">Accuracy</th>
<th align="left">Recall</th>
<th align="left">Precision</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="4">Case 1: COVID-19 <italic>vs.</italic> normal</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Max</td>
<td align="left" rowspan="2"><bold>0.975</bold></td>
<td align="left">0.988</td>
<td align="left">0.984</td>
<td align="left">0.975</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.984</td>
<td align="left">0.988</td>
<td align="left">0.975</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Avg</td>
<td align="left" rowspan="2">0.963</td>
<td align="left">0.970</td>
<td align="left">0.957</td>
<td align="left">0.963</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left">0.955</td>
<td align="left">0.970</td>
<td align="left">0.963</td>
</tr>
<tr>
<td align="left" rowspan="4">Case 2: COVID-19 <italic>vs.</italic><break/>non-COVID-19</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Max</td>
<td align="left" rowspan="2"><bold>0.981</bold></td>
<td align="left">0.967</td>
<td align="left">0.992</td>
<td align="left">0.974</td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.995</td>
<td align="left">0.980</td>
<td align="left">0.985</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Avg</td>
<td align="left" rowspan="2">0.963</td>
<td align="left">0.925</td>
<td align="left">0.974</td>
<td align="left">0.949</td>
</tr>
<tr>
<td align="left">Non-COVID-19</td>
<td align="left">0.985</td>
<td align="left">0.958</td>
<td align="left">0.971</td>
</tr>
<tr>
<td align="left" rowspan="6">Case 3: COVID-19 <italic>vs.</italic> normal <italic>vs.</italic> others</td>
<td align="left">Normal</td>
<td align="left" rowspan="3">Max</td>
<td align="left">0.772</td>
<td align="left">0.740</td>
<td align="left">0.607</td>
<td align="left">0.658</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.977</bold></td>
<td align="left">0.987</td>
<td align="left">0.948</td>
<td align="left">0.963</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.774</td>
<td align="left">0.632</td>
<td align="left">0.760</td>
<td align="left">0.687</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left" rowspan="3">Avg</td>
<td align="left">0.764</td>
<td align="left">0.704</td>
<td align="left">0.595</td>
<td align="left">0.645</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left">0.957</td>
<td align="left">0.972</td>
<td align="left">0.900</td>
<td align="left">0.934</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="left">0.763</td>
<td align="left">0.595</td>
<td align="left">0.750</td>
<td align="left">0.663</td>
</tr>
<tr>
<td align="left" rowspan="10">Case 4: 5-class classification</td>
<td align="left">Normal</td>
<td align="left" rowspan="5">Max</td>
<td align="left">0.924</td>
<td align="left">0.690</td>
<td align="left">0.976</td>
<td align="left">0.784</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.978</bold></td>
<td align="left">0.960</td>
<td align="left">0.937</td>
<td align="left">0.947</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.990</bold></td>
<td align="left">0.990</td>
<td align="left">0.973</td>
<td align="left">0.976</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.839</td>
<td align="left">0.640</td>
<td align="left">0.592</td>
<td align="left">0.610</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.857</td>
<td align="left">0.730</td>
<td align="left">0.630</td>
<td align="left">0.661</td>
</tr>
<tr>
<td align="left">Normal</td>
<td align="left" rowspan="5">Avg</td>
<td align="left">0.901</td>
<td align="left">0.553</td>
<td align="left">0.922</td>
<td align="left">0.684</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left"><bold>0.964</bold></td>
<td align="left">0.951</td>
<td align="left">0.883</td>
<td align="left">0.915</td>
</tr>
<tr>
<td align="left">Pneumonia</td>
<td align="left"><bold>0.984</bold></td>
<td align="left">0.972</td>
<td align="left">0.949</td>
<td align="left">0.960</td>
</tr>
<tr>
<td align="left">Pneumothorax</td>
<td align="left">0.816</td>
<td align="left">0.620</td>
<td align="left">0.538</td>
<td align="left">0.575</td>
</tr>
<tr>
<td align="left">Atelectasis</td>
<td align="left">0.842</td>
<td align="left">0.674</td>
<td align="left">0.594</td>
<td align="left">0.630</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="table-10">Tab. 10</xref>, the SVM model gave very high accuracy and overall performance in all experimental cases, with 0.99 accuracy for COVID-19 classification, while less accuracy was obtained for normal or other disease classification. Other measures which are recall, precision and F1-score were obtained with similar results as the accuracy as shown in the table.</p>
<p><xref ref-type="table" rid="table-11">Tab. 11</xref> presented the results obtained with the RF model, where COVID-19 detection was obtained with 0.96&#x2013;0.98 accuracy and similar values for recall, precision, and F1-score. The RF is not as good as the SVM in solving these classification problems.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Computation Time</title>
<p>In our experiments, deep learning and classical machine learning models were trained to classify the X-ray images. We wanted to compare our computation time between using deep learning (VGG19) and classical machine learning (SVM and RF). The computation time is separated into two parts which are image pre-processing time and training time. The pre-processing time includes the time when images are resized, reformatted, encoded, and feature extracted. This is shown in <xref ref-type="table" rid="table-12">Tab. 12</xref>. The computation time used for model training is shown in <xref ref-type="table" rid="table-13">Tab. 13</xref>.</p>
<table-wrap id="table-12"><label>Table 12</label>
<caption>
<title>Image pre-processing time (s)</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Models</th>
<th align="left">Case 1</th>
<th align="left">Case 2</th>
<th align="left">Case 3</th>
<th align="left">Case 4</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Deep learning: VGG 16, VGG19</td>
<td align="left">9</td>
<td align="left">11</td>
<td align="left">35</td>
<td align="left">33</td>
</tr>
<tr>
<td align="left">Classical Machine Learning with VGG16 features</td>
<td align="left">67</td>
<td align="left">77</td>
<td align="left">90</td>
<td align="left">62</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-13"><label>Table 13</label>
<caption>
<title>Computation time for model training (s)</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">fold/holdout</th>
<th align="left">Case 1</th>
<th align="left">Case 2</th>
<th align="left">Case 3</th>
<th align="left">Case 4</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">VGG16</td>
<td align="left">hold out</td>
<td align="left">2920</td>
<td align="left">4384</td>
<td align="left">1943</td>
<td align="left">1207</td>
</tr>
<tr>
<td align="left">VGG19</td>
<td align="left">hold out</td>
<td align="left">3522</td>
<td align="left">4659</td>
<td align="left">2328</td>
<td align="left">1860</td>
</tr>
<tr>
<td align="left" rowspan="6">SVM</td>
<td align="left">fold 1</td>
<td align="left">294</td>
<td align="left">507</td>
<td align="left">1955</td>
<td align="left">752</td>
</tr>
<tr>
<td align="left">fold 2</td>
<td align="left">284</td>
<td align="left">496</td>
<td align="left">1904</td>
<td align="left">751</td>
</tr>
<tr>
<td align="left">fold 3</td>
<td align="left">290</td>
<td align="left">499</td>
<td align="left">1875</td>
<td align="left">740</td>
</tr>
<tr>
<td align="left">fold 4</td>
<td align="left">286</td>
<td align="left">504</td>
<td align="left">1790</td>
<td align="left">726</td>
</tr>
<tr>
<td align="left">fold 5</td>
<td align="left">278</td>
<td align="left">485</td>
<td align="left">1832</td>
<td align="left">715</td>
</tr>
<tr>
<td align="left">average</td>
<td align="left">286.4</td>
<td align="left">498.2</td>
<td align="left">1871.2</td>
<td align="left">736.8</td>
</tr>
<tr>
<td align="left" rowspan="6">RF</td>
<td align="left">fold 1</td>
<td align="left">42</td>
<td align="left">58</td>
<td align="left">109</td>
<td align="left">63</td>
</tr>
<tr>
<td align="left">fold 2</td>
<td align="left">41</td>
<td align="left">57</td>
<td align="left">110</td>
<td align="left">61</td>
</tr>
<tr>
<td align="left">fold 3</td>
<td align="left">42</td>
<td align="left">58</td>
<td align="left">110</td>
<td align="left">62</td>
</tr>
<tr>
<td align="left">fold 4</td>
<td align="left">42</td>
<td align="left">59</td>
<td align="left">115</td>
<td align="left">63</td>
</tr>
<tr>
<td align="left">fold 5</td>
<td align="left">42</td>
<td align="left">59</td>
<td align="left">152</td>
<td align="left">64</td>
</tr>
<tr>
<td align="left">average</td>
<td align="left">41.8</td>
<td align="left">58.2</td>
<td align="left">119.2</td>
<td align="left">62.6</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The image pre-processing time for the deep learning model is less than the time required for image pre-processing for classical machine learning because the deep learning model does not require feature extraction. However, training each deep learning model takes a significantly longer time than training a classical machine learning model, as shown in <xref ref-type="table" rid="table-13">Tab. 13</xref>. In addition, for experimental case 3, the amount of data used in training each model is a lot more than in other cases. Thus, the SVM used more time with an average of 1871.2 s which is close to the deep learning training time. However, RF is the model that requires the lowest training time in all experimental cases.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Exploring NIH Dataset for COVID-19 Detection</title>
<p>To further demonstrate the effectiveness of our SVM model to identify COVID-19 infected images apart from many other lung-infected/abnormal and normal images, we consider a well-known NIH dataset that consists of X-ray images of many lung diseases and abnormalities. The dataset does not have COVID-19 X-ray images, so we include in this last experiment the set of COVID-19 images that were used in our previous experimental cases. So, this last experiment consists of 3500 COVID-19 images and 3500 non-COVID-19 images from the NIH database.</p>
<p>We sampled 3500 images from the dataset consisting of normal (no finding), multi-finding, and many lung diseases and abnormalities as shown in <xref ref-type="table" rid="table-14">Tab. 14</xref>. We want to perform COVID-19 detection against other diseases as well as normal. This is a binary class classification with 5-fold cross-validation. In this case, the VGG19 pre-trained features are used as input and the classification result is shown in <xref ref-type="table" rid="table-15">Tab. 15</xref>.</p>
<table-wrap id="table-14"><label>Table 14</label>
<caption>
<title>Sample dataset from NIH group as a non-COVID-19 class</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Finding</th>
<th align="left">Amount (images)</th>
<th align="left">Finding</th>
<th align="left">Amount (images)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">No finding</td>
<td align="left">2054</td>
<td align="left">Pleural Thickening</td>
<td align="left">49</td>
</tr>
<tr>
<td align="left"><bold>Multi finding</bold></td>
<td align="left">504</td>
<td align="left">Cardiomegaly</td>
<td align="left">37</td>
</tr>
<tr>
<td align="left">Infiltration</td>
<td align="left">273</td>
<td align="left">Fibrosis</td>
<td align="left">36</td>
</tr>
<tr>
<td align="left"><bold>Atelectasis</bold></td>
<td align="left">117</td>
<td align="left">Emphysema</td>
<td align="left">32</td>
</tr>
<tr>
<td align="left">Effusion</td>
<td align="left">110</td>
<td align="left">Consolidation</td>
<td align="left">17</td>
</tr>
<tr>
<td align="left">Nodule</td>
<td align="left">107</td>
<td align="left"><bold>Pneumonia</bold></td>
<td align="left">5</td>
</tr>
<tr>
<td align="left"><bold>Pneumothorax</bold></td>
<td align="left">89</td>
<td align="left">Hernia</td>
<td align="left">4</td>
</tr>
<tr>
<td align="left">Mass</td>
<td align="left">65</td>
<td align="left">Edema</td>
<td align="left">1</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-15"><label>Table 15</label>
<caption>
<title>Performance on non-COVID-19 and COVID-19 classes using SVM with VGG19 pre-trained features</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left" rowspan="2">Experiment</th>
<th align="left" rowspan="2">Classes</th>
<th align="left" rowspan="2">Stats</th>
<th align="center" colspan="4">Performance Metrics</th>
</tr>
<tr>
<th align="left">Accuracy</th>
<th align="left">Recall</th>
<th align="left">Precision</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" rowspan="4">COVID-19 <italic>vs.</italic><break/>non-COVID-19</td>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Max</td>
<td align="left" rowspan="2"><bold>0.992</bold></td>
<td align="left"><bold>0.997</bold></td>
<td align="left">0.992</td>
<td align="left"><bold>0.992</bold></td>
</tr>
<tr>
<td align="left">non-COVID-19</td>
<td align="left">0.992</td>
<td align="left"><bold>0.996</bold></td>
<td align="left">0.992</td>
</tr>
<tr>
<td align="left">COVID-19</td>
<td align="left" rowspan="2">Avg</td>
<td align="left" rowspan="2">0.984</td>
<td align="left"><bold>0.990</bold></td>
<td align="left">0.980</td>
<td align="left"><bold>0.985</bold></td>
</tr>
<tr>
<td align="left">non-COVID-19</td>
<td align="left">0.979</td>
<td align="left"><bold>0.990</bold></td>
<td align="left">0.984</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this experiment with our SVM model, COVID-19 images can be detected with 0.984 accuracy on average of 5 folds, and sometimes as high as 0.992 accuracy in a fold which is quite impressive. The other performance metrics which are recall, precision and F1-score also have very high values ranging between 0.98&#x2013;0.99 on average, and sometimes as high as 0.997, as shown in <xref ref-type="table" rid="table-15">Tab. 15</xref>. This means the SVM model can detect COVID-19 accurately against many other findings in the chest X-ray images.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Summary and Conclusion</title>
<p>In summary, our classification models have very good performance on binary-class classification with both classical machine learning approach and deep learning approach. Overall, the classical machine learning technique SVM with VGG-19 pre-trained features gives the best classification performance. In this work, we also consider the computation time of all classification models including image pre-processing time. Datasets with many lung diseases and abnormalities from well-known database and NIH were also considered in our classification models.</p>
<p>The SVM model requires more computation time (image pre-processing time and training time) than the RF model, but less than the deep learning&#x2019;s computation time. In other words, the proposed SVM with VGG19 pre-trained features has superior performance than the deep learning models both in terms of classification performance (i.e., accuracy, recall, precision, and F1-score) and computation time.</p>
<p>In conclusion, we can apply deep learning and classical machine learning approaches to automatically detect/classify lung diseases with chest X-ray images. Our future work is to deploy the detection/classification models for practical usage.</p>
</sec>
</body>
<back><fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="web"><person-group person-group-type="author"><collab>WHO Coronavirus (Covid-19)</collab></person-group>. [Online]. Available: <uri xlink:href="https://covid19.who.int/">https://covid19.who.int/</uri> (20 October <year>2021</year>).</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. D.</given-names> <surname>Arias-Londo&#x00F1;o</surname></string-name>, <string-name><given-names>J. A.</given-names> <surname>G&#x00F3;mez-Garc&#x00ED;a</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Moro-Vel&#x00E1;zquez</surname></string-name> and <string-name><given-names>J. I.</given-names> <surname>Godino-Lorente</surname></string-name></person-group>, &#x201C;<article-title>Artificial intelligence applied to chest X-ray images for the automatic detection of COVID-19. a thoughtful evaluation approach</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>226811</fpage>&#x2013;<lpage>226827</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. I.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>J. L.</given-names> <surname>Shah</surname></string-name> and <string-name><given-names>M. M.</given-names> <surname>Bhat</surname></string-name></person-group>, &#x201C;<article-title>CoroNet: A deep neural network for detection and diagnosis of COVID-19 from chest x-ray images</article-title>,&#x201D; <source>Computer Methods and Programs in Biomedicine</source>, vol. <volume>196</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. N.</given-names> <surname>Vinod</surname></string-name> and <string-name><given-names>S. R. S.</given-names> <surname>Prabaharan</surname></string-name></person-group>, &#x201C;<article-title>Data science and the role of artificial intelligence in achieving the fast diagnosis of COVID-19</article-title>,&#x201D; <source>Chaos, Solitons &#x0026; Fractals</source>, vol. <volume>140</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Abbas</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Abdelsamea</surname></string-name> and <string-name><given-names>M. M.</given-names> <surname>Gaber</surname></string-name></person-group>, &#x201C;<article-title>Classification of COVID-19 in chest X-ray images using deTraC deep convolutional neural network</article-title>,&#x201D; <source>Applied Intelligence</source>, vol. <volume>51</volume>, pp. <fpage>854</fpage>&#x2013;<lpage>864</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sedik</surname></string-name>, <string-name><given-names>A. M.</given-names> <surname>Iliyasu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Abd El-Rahiem</surname></string-name>, <string-name><given-names>M. E.</given-names> <surname>Abdel Samea</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Abdel-Raheem</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Deploying machine and deep learning models for efficient data-augmented detection of COVID-19 infections</article-title>,&#x201D; <source>Viruses</source>, vol. <volume>12</volume>, no. <issue>7</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>29</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. Y.</given-names> <surname>Kamil</surname></string-name></person-group>, &#x201C;<article-title>A deep learning framework to detect covid-19 disease via chest X-ray and CT scan images</article-title>,&#x201D; <source>International Journal of Electrical and Computer Engineering</source>, vol. <volume>11</volume>, pp. <fpage>844</fpage>&#x2013;<lpage>850</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. K.</given-names> <surname>Rangarajan</surname></string-name> and <string-name><given-names>H. K.</given-names> <surname>Ramachandran</surname></string-name></person-group>, &#x201C;<article-title>A preliminary analysis of AI based smartphone application for diagnosis of COVID-19 using chest X-ray images</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>183</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I. J.</given-names> <surname>Goodfellow</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Pouget-Abadie</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Mirza</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Warde-Farley</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Generative adversarial networks</article-title>,&#x201D; in <conf-name>Int. Conf. on Neural Information Processing Systems</conf-name>, <conf-loc>Montreal, Canada</conf-loc>, pp. <fpage>2672</fpage>&#x2013;<lpage>2680</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. I.</given-names> <surname>Mor&#x00ED;s</surname></string-name>, <string-name><given-names>J. J.</given-names> <surname>de Moura Ramos</surname></string-name>, <string-name><given-names>J. N.</given-names> <surname>Buj&#x00E1;n</surname></string-name> and <string-name><given-names>M. O.</given-names> <surname>Hortas</surname></string-name></person-group>, &#x201C;<article-title>Data augmentation approaches using cycle-consistent adversarial networks for improving COVID-19 screening in portable chest X-ray images</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>185</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. B.</given-names> <surname>Chandra</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Verma</surname></string-name>, <string-name><given-names>B. K.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Jain</surname></string-name> and <string-name><given-names>S. S.</given-names> <surname>Netam</surname></string-name></person-group>, &#x201C;<article-title>Coronavirus disease (COVID-19) detection in chest X-ray images using majority voting based classifier ensemble</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>165</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Emary</surname></string-name>, <string-name><given-names>H. M.</given-names> <surname>Zawbaa</surname></string-name> and <string-name><given-names>A. E.</given-names> <surname>Hassanien</surname></string-name></person-group>, &#x201C;<article-title>Binary grey wolf optimization approaches for feature selection</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>172</volume>, pp. <fpage>371</fpage>&#x2013;<lpage>381</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Tuncer</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dogan</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Ozyurt</surname></string-name></person-group>, &#x201C;<article-title>An automated residual exemplar local binary pattern and iterative ReliefF based COVID-19 detection method using chest X-ray image</article-title>,&#x201D; <source>Chemometrics and Intelligent Laboratory Systems</source>, vol. <volume>203</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Pietik&#x00E4;inen</surname></string-name></person-group>, &#x201C;<article-title>Image analysis with local binary patterns</article-title>,&#x201D; in <conf-name>Image Analysis</conf-name>, <conf-loc>Joensuu, Finland</conf-loc>, pp. <fpage>115</fpage>&#x2013;<lpage>118</lpage>, <year>2005</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. J.</given-names> <surname>Urbanowicz</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Meeker</surname></string-name>, <string-name><given-names>W.</given-names> <surname>LaCava</surname></string-name>, <string-name><given-names>R. S.</given-names> <surname>Olson</surname></string-name> and <string-name><given-names>J. H.</given-names> <surname>Moore</surname></string-name></person-group>, &#x201C;<article-title>Relief-based feature selection: Introduction and review</article-title>,&#x201D; <source>Journal of Biomedical Informatics</source>, vol. <volume>85</volume>, pp. <fpage>189</fpage>&#x2013;<lpage>203</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>&#x015E;</given-names> <surname>&#x00D6;zt&#x00FC;rk</surname></string-name>, <string-name><given-names>U.</given-names> <surname>&#x00D6;zkaya</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Barstu&#x011F;an</surname></string-name></person-group>, &#x201C;<article-title>Classification of coronavirus (COVID-19) from X-ray and CT images using shrunken features</article-title>,&#x201D; <source>International Journal of Imaging Systems and Technology</source>, vol. <volume>31</volume>, pp. <fpage>5</fpage>&#x2013;<lpage>15</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. A.</given-names> <surname>Ardakani</surname></string-name>, <string-name><given-names>U. R.</given-names> <surname>Acharya</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Habibollahi</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Mohammadi</surname></string-name></person-group>, &#x201C;<article-title>COVIDiag: A clinical CAD system to diagnose COVID-19 pneumonia based on CT findings</article-title>,&#x201D; <source>European Radiology</source>, vol. <volume>31</volume>, pp. <fpage>121</fpage>&#x2013;<lpage>130</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. A. D.</given-names> <surname>J&#x00FA;nior</surname></string-name>, <string-name><given-names>L. B.</given-names> <surname>da Cruz</surname></string-name>, <string-name><given-names>J. O. B.</given-names> <surname>Diniz</surname></string-name>, <string-name><given-names>G. L. F.</given-names> <surname>da Silva</surname></string-name>, <string-name><given-names>G. B.</given-names> <surname>Junior</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Automatic method for classifying COVID-19 patients based on chest X-ray images, using deep features and PSO-optimized XGBoost</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>183</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Tamal</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Alshammari</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Alabdullah</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Hourani</surname></string-name>, <string-name><given-names>H. A.</given-names> <surname>Alola</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>An integrated framework with machine learning and radiomics for accurate and rapid early diagnosis of COVID-19 from chest X-ray</article-title>,&#x201D; <source>Expert Systems with Applications</source>, vol. <volume>180</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Lambin</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Rios-Velazquez</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Leijenaar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Carvalho</surname></string-name>, <string-name><given-names>R. G. P. M.</given-names> <surname>Van Stiphout</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Radiomics: Extracting more information from medical images using advanced feature analysis</article-title>,&#x201D; <source>European Journal of Cancer</source>, vol. <volume>48</volume>, pp. <fpage>441</fpage>&#x2013;<lpage>446</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. E. H.</given-names> <surname>Chowdhury</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Khandakar</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Mazhar</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Kadir</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Can AI help in screening viral and COVID-19 pneumonia?</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>132665</fpage>&#x2013;<lpage>132676</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Khandakar</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Qiblawey</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Tahir</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kiranyaz</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Exploring the effect of image enhancement techniques on COVID-19 detection using chest X-ray images</article-title>,&#x201D; <source>Computers in Biology and Medicine</source>, vol. <volume>132</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bagheri</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>ChestX-Ray8: Hospital-scale chest X-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases</article-title>,&#x201D; in <conf-name>IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Hawaii, Honolulu USA</conf-loc>, pp. <fpage>3462</fpage>&#x2013;<lpage>3471</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Krizhevsky</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Sutskever</surname></string-name> and <string-name><given-names>G. E.</given-names> <surname>Hinton</surname></string-name></person-group>, &#x201C;<article-title>ImageNet classification with deep convolutional neural networks</article-title>,&#x201D; <source>Communications of the ACM</source>, vol. <volume>60</volume>, pp. <fpage>84</fpage>&#x2013;<lpage>90</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. J.</given-names> <surname>Pan</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>A survey on transfer learning</article-title>,&#x201D; <source>IEEE Transactions on Knowledge and Data Engineering</source>, vol. <volume>22</volume>, pp. <fpage>1345</fpage>&#x2013;<lpage>1359</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Socher</surname></string-name>, <string-name><given-names>L. -J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Li</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>ImageNet: A large-scale hierarchical image database</article-title>,&#x201D; in <conf-name>IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Florida, Miami, USA</conf-loc>, pp. <fpage>248</fpage>&#x2013;<lpage>255</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Simonyan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Very deep convolutional networks for large-scale image recognition</article-title>,&#x201D; in <conf-name>Int. Conf. on Learning Representations</conf-name>, <conf-loc>California, San Diego, USA</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Hearst</surname></string-name>, <string-name><given-names>S. T.</given-names> <surname>Dumais</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Osuna</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Platt</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Scholkopf</surname></string-name></person-group>, &#x201C;<article-title>Support vector machines</article-title>,&#x201D; <source>IEEE Intelligent Systems and Their Applications</source>, vol. <volume>13</volume>, no. <issue>4</issue>, pp. <fpage>18</fpage>&#x2013;<lpage>28</lpage>, <year>1998</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Breiman</surname></string-name></person-group>, &#x201C;<article-title>Random forests</article-title>,&#x201D; <source>Machine Learning</source>, vol. <volume>45</volume>, pp. <fpage>5</fpage>&#x2013;<lpage>32</lpage>, <year>2001</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. R.</given-names> <surname>Quinlan</surname></string-name></person-group>, &#x201C;<article-title>Induction of decision trees</article-title>,&#x201D; <source>Machine Learning</source>, vol. <volume>1</volume>, pp. <fpage>81</fpage>&#x2013;<lpage>106</lpage>, <year>1986</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>