<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">29663</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2022.029663</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Face Mask Recognition for Covid-19 Prevention</article-title>
<alt-title alt-title-type="left-running-head">Face Mask Recognition for Covid-19 Prevention</alt-title>
<alt-title alt-title-type="right-running-head">Face Mask Recognition for Covid-19 Prevention</alt-title>
</title-group>
<contrib-group content-type="authors"><contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Luu</surname><given-names>Trong Hieu</given-names>
</name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Phuc</surname><given-names>Phan Nguyen Ky</given-names>
</name><xref ref-type="aff" rid="aff-2">2</xref><email>pnkphuc@hcmiu.edu.vn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Yu</surname><given-names>Zhiqiu</given-names>
</name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Pham</surname><given-names>Duy Dung</given-names>
</name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Cao</surname><given-names>Huu Trong</given-names>
</name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>College of Engineering, Can Tho University</institution>, <addr-line>Can Tho city, 910900</addr-line>, <country>Viet Nam</country></aff>
<aff id="aff-2"><label>2</label><institution>International University-Vietnam National University, Vietnam National University</institution>, <addr-line>HoChiMinh City, 70000</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Industrial Management, National Taiwan University of Science and Technology</institution>, <addr-line>Taipei, 10607</addr-line>, <country>Taiwan</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Phan Nguyen Ky Phuc. Email: <email>pnkphuc@hcmiu.edu.vn</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-06-14"><day>14</day>
<month>06</month>
<year>2022</year></pub-date>
<volume>73</volume>
<issue>2</issue>
<fpage>3251</fpage>
<lpage>3262</lpage>
<history>
<date date-type="received">
<day>09</day>
<month>3</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>5</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Luu et al.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Luu et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_29663.pdf"></self-uri>
<abstract>
<p>In recent years, the COVID-19 pandemic has negatively impacted all aspects of social life. Due to ease in the infected method, i.e., through small liquid particles from the mouth or the nose when people cough, sneeze, speak, sing, or breathe, the virus can quickly spread and create severe problems for people&#x2019;s health. According to some research as well as World Health Organization (WHO) recommendation, one of the most economical and effective methods to prevent the spread of the pandemic is to ask people to wear the face mask in the public space. A face mask will help prevent the droplet and aerosol from person to person to reduce the risk of virus infection. This simple method can reduce up to 95% of the spread of the particles. However, this solution depends heavily on social consciousness, which is sometimes unstable. In order to improve the effectiveness of wearing face masks in public spaces, this research proposes an approach for detecting and warning a person who does not wear or misuse the face mask. The approach uses the deep learning technique that relies on GoogleNet, AlexNet, and VGG16 models. The results are synthesized by an ensemble method, i.e., the bagging technique. From the experimental results, the approach represents a more than 95% accuracy of face mask recognition.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Face mask detection</kwd>
<kwd>deep learning</kwd>
<kwd>AlexNet</kwd>
<kwd>GoogLeNet</kwd>
<kwd>VGG16</kwd>
<kwd>ensemble method</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>According to World Health Organization (WHO), Coronavirus disease (COVID-19) is an infectious disease caused by the SARS-CoV-2 virus. Depending on medical conditions, infected people suffer different types of symptoms. In most cases, people undergo mild to moderate respiratory illness and recover without requiring special treatment. However, due to the cytokine storm, i.e., excessive production of cytokines, some people have experienced severe aggravation and widespread tissue damage, which cause multi-organ failure and can lead to death. Older infected people and those with underlying medical conditions like cardiovascular disease, diabetes, chronic respiratory disease, or cancer are more likely to develop severe illness and often require special treatments. Anyone can get sick with COVID-19 and become seriously ill or die at any age. According to the data of WHO, up to the time of this study, there are more than 420 million patients, and nearly 6 million people have passed away [<xref ref-type="bibr" rid="ref-1">1</xref>]. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows the statistics of COVID-19 confirmed cases from December 2019 to December 2021.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The statistics of COVID-19 confirmed cases from December 2019 to December 2021</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-1.png"/>
</fig>
<p>According to <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the death number had remained stable and had not decreased until the end of December 2021, when the pandemic lasted for more than two years and several treatments were proposed.</p>
<p>Apart from the high fatality rate, another prominent risk of this pandemic is that the patients can also continue to experience severe symptoms after their initial recovery even after receiving proper treatments. These health issues are also sometimes called &#x201C;post-COVID-19 conditions&#x201D; [<xref ref-type="bibr" rid="ref-2">2</xref>]. These are not limited to only elders and people with many severe medical conditions, and even young and healthy people can feel unwell for weeks to months after infection. Some common symptoms of &#x201C;post-COVID-19 conditions&#x201D; include fatigue, shortness of breath or difficulty breathing, cough, joint pain, chest pain, etc. Most of these symptoms derive from organ damage under the virus&#x2019;s effects. Although several types of vaccines have been developed and deployed in numerous countries, the appearance of new variants of COVID-19 and the attenuation of vaccine effects over time also create new challenges to slow down and prevent the spread of the pandemic. Currently, the most effective and cheapest method to protect people from infection is to stay at least one meter apart from others, wear a properly fitted mask, and wash hands with an alcohol-based rub frequently.</p>
<p>The major infected way of the virus is through small liquid particles from the mouth or nose when people cough, sneeze, speak, sing, or breathe. These particles range from larger respiratory droplets to smaller aerosols. Due to the infected mechanism, wearing the face mask properly can reduce up to 95% of the spread of the particles [<xref ref-type="bibr" rid="ref-3">3</xref>]. Recognizing the importance of properly face mask usage against COVID-19, this study tries to develop an effective method to classify whether a person has used the face mask correctly in the public area. By detecting people without a face mask or inappropriately wearing a face mask, the system can give a warning and show the image on the large screen so that other people can realize the potential risk and keep their distance from them. These systems are extremely helpful in the public space with a high density of people, such as supermarkets, classrooms, restaurants, or waiting rooms of airports. The system will save the headcounts as well as increase the consciousness of people of how the importance of wearing the face mask properly in the public area.</p>
<p>The rest of this study is organized as follows. Works related to our approach to deep learning and convolutional neural networks are reviewed in Section 2. In addition, the proposed methodology of this study is also described in this section. The experimental setup and results are presented in Section 3, while the conclusion and future works are presented in Section 4.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works &#x0026; Proposed Method</title>
<p>Convolution neural networks (CNN), first introduced by LeCun [<xref ref-type="bibr" rid="ref-4">4</xref>], is a deep learning model which can obtain high accuracy for the classification task. CNN principle is based on activities of the visual perception of the human brain when recognizing simple shapes. In general, CNN is constructed by using three different primal layer types, i.e., convolution layers, nonlinear layers, and filtering layers, which are stacked in several sequential orders, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The CNN structure proposed by LeCun</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-2.png"/>
</fig>
<p>In the proposed network, each layer is responsible for a specific function in the classification process. For example, the convolution layers will apply different filters to extract abstract features in the image content. A rectifier linear unit layer is often put behind the convolutional layer to produce the nonlinear property as well as more abstract information for subsequent layers. Apart from convolutional layers and rectifier linear unit layers, pooling layers are often employed in CNN to reduce the sample size while still retaining the most promising features for classifications. Lecun&#x2019;s study motivated other researchers and then created an explosion in introducing new deep learning models for image classification. Three of the most successful models adopted in this study are AlexNet, GoogLeNet, and VGG16. The structure of the AlexNet is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<p>The AlexNet network was first introduced in [<xref ref-type="bibr" rid="ref-5">5</xref>]. This network won the first on The ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2012. The AlexNet comprises several layers, including an input layer, five convolutional layers, five ReLu layers, two normalization layers, three pooling layers, three fully connected layers, a drop layer, and a softmax output layer. AlexNet takes an RGB image with the size of 224 &#x00D7; 224 as an input, and the output is a vector with the size of 1000 &#x00D7; 1. The prominent point in the AlexNet model is the employment of the dropout technique to prevent overfitting. In this technique, each node has the probability p to be selected for training in each epoch.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The AlexNet structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-3.png"/>
</fig>
<p>Besides the AlexNet, this study also adopts the GoogLeNet [<xref ref-type="bibr" rid="ref-6">6</xref>] for the classification task. This network resulted from the combination between Google company and Cornell University, and it also won the ILSVRC in 2014. The structure of GoogLeNet is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The GoogLeNet structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-4.png"/>
</fig>
<p>The GoogLeNet differentiates itself from other deep learning models by using nine inceptions blocks from 22 deep layers and five pooling layers. These blocks allow the networks to conduct parallel learning. In other words, one input, whose size is 224 &#x00D7; 224, can be fed to several different convolutional layers to create different output features. These features then can be concatenated into an output. The parallel learning process helps the network itself to study and obtain more features than the traditional CNN approach. Moreover, the network also applies the 1 &#x00D7; 1 convolutional block to decrease the network size so as to quickly train the network. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows the structure of an inception block in GoogLeNet. In the inception block, the outputs from the former layer are used as inputs, then fed into four different branches, and finally concatenated into one output.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The intercept block structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-5.png"/>
</fig>
<p>The last deep neural network integrated into the proposed model is VGG16. VGG16 [<xref ref-type="bibr" rid="ref-7">7</xref>] is a Convolutional Neural Network architecture first introduced in 2014 and achieved 92.7% top-5 test accuracy in ImageNet. The structure of VGG16 is described below in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The VGG16 structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-6.png"/>
</fig>
<p>VGG16 takes an RGM image with size 224 &#x00D7; 224 as an input. This input is passed to the whole network then the resulting output is passed to other layers to automatically extract features and make the classification. It is noted that the fourteen and fifteen layers are fully connected hidden layers of 4096 units, followed by a softmax output layer of 1000 units.</p>
<p>To utilize the power of these successful networks as well as to reduce the training time of these models, transfer learning and ensemble technique are applied in later steps. Generally, the main objective of transfer learning is to store knowledge gained while solving one problem and using it for a different but related problem [<xref ref-type="bibr" rid="ref-8">8</xref>]. In the case of CNN, transfer learning often relates to retraining some layers of well-trained models while freezing other layers. The selected layers for retraining are often the last layers of the model where the softmax function is applied to conduct classification. In this study, transfer learning is applied to all networks above.</p>
<p>In order to create the proper inputs for this network, the Viola-Jones algorithm [<xref ref-type="bibr" rid="ref-9">9</xref>] is applied to detect and extract faces from the image. The Viola-Jones algorithm basically comprises four main operations, including applying a Haar-like filter, calculating the integrated image, applying the Ada Boost algorithm, and returning the results based on a cascading classifier. The principle of the cascade classifier is given in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Principle of cascade algorithm</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-7.png"/>
</fig>
<p>It is noted that in our proposed system, one raw image can have multiple faces that can belong to different label classes. As a result of that, the Viola-Jones algorithm returns several sub images of faces. Each face image must be passed to preprocessing phase for enhancing or downsizing so that it can be formatted properly, i.e., a size of 224 &#x00D7; 224, and sent to all networks as the inputs. Even though there are other methods for face and object detection [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>], this study still adopts the Viola-Jones algorithm due to its robust and fast computation.</p>
<p>The returned output vectors from each network are then recollected for ensembling. In this study, the bagging technique of the voting method is applied at the ensembling stage for making final decisions and final classifications [<xref ref-type="bibr" rid="ref-14">14</xref>]. The final label is decided by majority rule. For example, if output vectors of AlexNet, GoogLeNet, and VGG16 are <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.8</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.7</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.6</mml:mn><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> respectively, according to the majority rule the image is classified as class 2 due to the agreement of AlexNet and GoogLeNet. If three networks give three different output results, then the label with the highest average weight is selected. For instance, if <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.8</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></inline-formula> <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.7</mml:mn><mml:mo>,</mml:mo><mml:mn>0.1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.2</mml:mn><mml:mo>,</mml:mo><mml:mn>0.6</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mover><mml:mi>v</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mn>3</mml:mn></mml:mfrac><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.4</mml:mn><mml:mo>,</mml:mo><mml:mn>0.33</mml:mn><mml:mo>,</mml:mo><mml:mn>0.27</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, the face is classified as class 1. The main use of the ensemble technique is to improve the overall performance of the entire system. By combining several independent weak classifiers, the ensemble method can create a strong classifier with higher sensitivity. The diagram of the proposed system is described in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>The proposed system structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-8.png"/>
</fig>
</sec>
<sec id="s3">
<label>3</label>
<title>Data &#x0026; Experimental Setup</title>
<p>The dataset, which was created by our team, includes three classes regarding three respective outputs as described above. The number of samples for each class label in the training set is shown in <xref ref-type="table" rid="table-1">Tab. 1</xref>. In the training set, several mask types as well as the number of faces in the image are selected to ensure diversity. However, one image only includes one type of output labels, as shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. Each image&#x2019;s resolution is 1280 &#x00D7; 960, and it will be resized to fit the input of each deep learning method. Only 80% in each class is used for training, and 20% is used for validation.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>The characteristic of the training data set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>Details</th>
<th>Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>C1</bold></td>
<td>No mask</td>
<td>1200</td>
</tr>
<tr>
<td><bold>C2</bold></td>
<td>Misuse</td>
<td>1200</td>
</tr>
<tr>
<td><bold>C3</bold></td>
<td>Wear mask</td>
<td>1200</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Images from the training data set</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-9.png"/>
</fig>
<p>In this study, the stochastic gradient descent method [<xref ref-type="bibr" rid="ref-15">15</xref>] is applied for training each network with a learning rate of approximately 0.0001. The batch size is 30, and the number of epochs is 120. The division of data set into small batches with the size of 30 is to ensure loading capability. Loading the whole training data set at once is impractical for most problems due to RAM limitations. In our cases, 80% of the training data set is 3600, which requires 120 epochs of training. The selected objective function is the cross-entropy loss function. In general, the learning rate must be chosen very carefully to avoid a low learning process as well as divergence. Overfitting is prevented by the dropout technique, and the pixel values of the input image are normalized between 0 and 1.</p>
<p>To test the accuracy of the proposed system in reality, the testing set is designed in a different way. In the testing set, each image can have more than one type of output labels. The testing set now includes nine classes. The details of the testing set are given in <xref ref-type="table" rid="table-2">Tab. 2</xref>, and images in the testing set are shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The details of the testing set</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>Details</th>
<th>Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>C1</bold></td>
<td>1 faces without mask</td>
<td>150</td>
</tr>
<tr>
<td><bold>C2</bold></td>
<td>2 faces without mask</td>
<td>150</td>
</tr>
<tr>
<td><bold>C3</bold></td>
<td>1 face with mask misuse</td>
<td>150</td>
</tr>
<tr>
<td><bold>C4</bold></td>
<td>2 faces with mask misuse</td>
<td>150</td>
</tr>
<tr>
<td><bold>C5</bold></td>
<td>1 face with mask</td>
<td>150</td>
</tr>
<tr>
<td><bold>C6</bold></td>
<td>2 faces with mask</td>
<td>150</td>
</tr>
<tr>
<td><bold>C7</bold></td>
<td>1 face with mask and 1 face with mask misuse</td>
<td>150</td>
</tr>
<tr>
<td><bold>C8</bold></td>
<td>1 face with mask and 1 face without mask</td>
<td>150</td>
</tr>
<tr>
<td><bold>C9</bold></td>
<td>1 face without mask and 1 face with mask misuse</td>
<td>150</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Images from the testing set</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-10.png"/>
</fig>
<p><xref ref-type="table" rid="table-3">Tab. 3</xref> presents the results of the proposed system in the form of a confusion matrix. Finally, <xref ref-type="table" rid="table-4">Tab. 4</xref> will summarize the performance of each network in the proposed system on the different classes of the test set. It is noted that, even though one class can be wrong classified as another class, the half of content inside the image of a class can still be accurate since the return output of the proposed system includes only &#x201C;NoMask&#x201D;, &#x201C;Misuse&#x201D;, and &#x201C;WearMask&#x201D;.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The confusion result matrix</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th rowspan="9">Actual label</th>
<th>C1</th>
<th>148</th>
<th>2</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
<th>&#x2013;</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>C2</bold></td>
<td>2</td>
<td>148</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C3</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>144</td>
<td>3</td>
<td>3</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C4</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>129</td>
<td>&#x2013;</td>
<td>8</td>
<td>13</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C5</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>149</td>
<td>1</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C6</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>3</td>
<td>147</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C7</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>21</td>
<td>&#x2013;</td>
<td>10</td>
<td>119</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>C8</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>141</td>
<td>9</td>
</tr>
<tr>
<td><bold>C9</bold></td>
<td>13</td>
<td>&#x2013;</td>
<td>6</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>131</td>
</tr>
<tr>
<td></td>
<td></td>
<td><bold>C1</bold></td>
<td><bold>C2</bold></td>
<td><bold>C3</bold></td>
<td><bold>C4</bold></td>
<td><bold>C5</bold></td>
<td><bold>C6</bold></td>
<td><bold>C7</bold></td>
<td><bold>C8</bold></td>
<td><bold>C9</bold></td>
</tr>
<tr>
<td></td>
<td></td>
<td colspan="9" align="center"><bold>Predicted label</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance of each network in the proposed system.</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Class</th>
<th>Total sample</th>
<th>Alexnet</th>
<th>Googlenet</th>
<th>VGG16</th>
<th>Sensitivity</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>C1</bold></td>
<td>150</td>
<td>98.6%</td>
<td>98.6%</td>
<td>98.6%</td>
<td>99.94%</td>
</tr>
<tr>
<td><bold>C2</bold></td>
<td>150</td>
<td>96.2%</td>
<td>97.4%</td>
<td>96%</td>
<td>99.65%</td>
</tr>
<tr>
<td><bold>C3</bold></td>
<td>150</td>
<td>94%</td>
<td>96.5%</td>
<td>94%</td>
<td>99.24%</td>
</tr>
<tr>
<td><bold>C4</bold></td>
<td>150</td>
<td>82.3%</td>
<td>86.1%</td>
<td>79.5%</td>
<td>92.07%</td>
</tr>
<tr>
<td><bold>C5</bold></td>
<td>150</td>
<td>98.6%</td>
<td>98.6%</td>
<td>98.6%</td>
<td>99.94%</td>
</tr>
<tr>
<td><bold>C6</bold></td>
<td>150</td>
<td>98.6%</td>
<td>98.6%</td>
<td>98.6%</td>
<td>99.94%</td>
</tr>
<tr>
<td><bold>C7</bold></td>
<td>150</td>
<td>77.2%</td>
<td>79.3%</td>
<td>76.1%</td>
<td>87.14%</td>
</tr>
<tr>
<td><bold>C8</bold></td>
<td>150</td>
<td>91%</td>
<td>94%</td>
<td>91%</td>
<td>98.21%</td>
</tr>
<tr>
<td><bold>C9</bold></td>
<td>150</td>
<td>83.7%</td>
<td>87.2%</td>
<td>81.1%</td>
<td>93.2%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>During the testing phase, we also test the capability of the proposed system with some very special types of face masks, i.e., the transparent face mask, and funny face mask which are shown in <xref ref-type="fig" rid="fig-11">Fig. 11</xref></p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Transparent and funny face masks</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_29663-fig-11.png"/>
</fig>
<p>Given the funny face masks, all three networks cannot perform really well. The proposed system cannot classify correctly in about 85% of test cases. This can be seen as the current limitations of our approach. In the case of transparent face masks, through carefully checking the output of each separate network, the probability that a network labeled it as the &#x201C;WearMask&#x201D; is not high compared to the other two classes. Each separate network often wrong assigned them to the label &#x201C;Misuse&#x201D; since both the mouth and nose also appear in the image while the rest of the face is covered by an exhalation valve. This phenomenon is the main explanation for the poor classification of class C7, which incidentally includes several images of transparent face masks. As explanations, these transparent face masks are often misclassified as &#x201C;Misuse&#x201D; with falls into C4.</p>
<p>It can be seen that, except in the case of transparent facemask and funny face mask, the proposed system has a very high sensitivity for the rest. It can be explained by the high sensitivity of each network. Given a class, if all networks are assumed as independent of other networks and the sensitivity of network <italic>i</italic> is <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, with three networks probability that network <italic>i</italic> correctly classified the given class as <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. There are four cases that the proposed system can correctly make the classification. The first case is that all three networks correctly make the classification. The rest is when one network assigns wrong labels, but the others make decisions correctly. The probability that here the sensitivity of proposed whole system when using voting process classifies a face correctly regarding probability theory is under this assumption is given by <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.</p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mi>y</mml:mi><mml:mi>s</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>It can be seen that this probability is much higher than the probability of one network.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Conclusion &#x0026; Future Works</title>
<p>Wearing face masks in an effective, simple, and cheap way can prevent the spread of the COVID virus in the public space. However, due to some reasons, not wearing a face mask and misusing a face mask still can be met very frequently in public spaces. These behaviors can spread the droplet and aerosol from person to person, which will increase the risk of virus infection. As a result of this, a system to detect such behavior and give proper warnings in public spaces is therefore required. This kind of system is extremely necessary for close public spaces such as university classes or offices. This study tries to propose a real time system to detect the proper usage of face masks in the public space. By combining Viola-Jones, three different deep learning neural networks, and the voting process, the proposed system has effectively handled normal face mask-wearing issues. However, the proposed system still has an obvious limitation on correctly labelling when it has to handle transparent and funny face masks. This limitation will be considered in future research.</p>
</sec>
</body>
<back>
<ack><p>The authors thank TopEdit (www.topeditsci.com) for its linguistic assistance during the preparation of this manuscript.</p></ack>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><article-title>Report of WHO on Coronavirus (COVID-19) Dashboard</article-title>. <uri>https://covid19.who.int/WHO-COVID-19-global-table-data.csv</uri>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. B.</given-names> <surname>Soriano</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Murthy</surname></string-name>, <string-name><given-names>J. C.</given-names> <surname>Marshall</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Relan</surname></string-name> and <string-name><given-names>J. V.</given-names> <surname>Diaz</surname></string-name></person-group>, &#x201C;<article-title>A clinical case definition of post-COVID-19 condition by a Delphi consensus</article-title>,&#x201D; <source>The Lancet Infectious Disease</source>, vol. <volume>1</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>A technical review of face mask wearing in preventing respiratory COVID-19 transmission</article-title>,&#x201D; <source>Current Opinion in Colloid &#x0026; Interface Science</source>, vol. <volume>52</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>19</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>LeCun</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Boser</surname></string-name>, <string-name><given-names>J. S.</given-names> <surname>Denker</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Henderson</surname></string-name>, <string-name><given-names>R. E.</given-names> <surname>Howard</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Backpropagation applied to handwritten zip code recognition</article-title>,&#x201D; <source>Neural Computation</source>, vol. <volume>1</volume>, no. <issue>4</issue>, pp. <fpage>541</fpage>&#x2013;<lpage>551</lpage>, <year>1989</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Krizhevsky</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Sutskever</surname></string-name> and <string-name><given-names>G. E.</given-names> <surname>Hinton</surname></string-name></person-group>, &#x201C;<article-title>ImageNet classification with deep convolution neural networks</article-title>,&#x201D; in <conf-name>Proc. NIPS 2012</conf-name>, <publisher-loc>Nevada, USA</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2012</year>. </mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Szegedy</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Jia</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Sermanet</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Reed</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Going deeper with convolutions</article-title>,&#x201D; <volume>1</volume>, <comment>arXiv.1409.4842</comment>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Simonyan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Very deep convolutional networks for large-scale image recognition</article-title>,&#x201D; in <conf-name>Proc. ICLR</conf-name>, <publisher-loc>San Diego, CA, USA</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2015</year>. </mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Weiss</surname></string-name>, <string-name><given-names>T. M.</given-names> <surname>Khoshgoftaar</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>A survey of transfer learning</article-title>,&#x201D; <source>Journal of Big Data</source>, vol. <volume>3</volume>, no. <issue>9</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>40</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Viola</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Jones</surname></string-name></person-group>, &#x201C;<article-title>Robust real-time face detection</article-title>,&#x201D; <source>International Journal of Computer Vision</source>, vol. <volume>57</volume>, no. <issue>10</issue>, pp. <fpage>137</fpage>&#x2013;<lpage>154</lpage>, <year>2004</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lei</surname></string-name> and <string-name><given-names>S. Z.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Aggregate channel features for multi-view face detection</article-title>,&#x201D; in <conf-name>Proc. IJCB</conf-name>, <publisher-loc>Clearwater, USA</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2014</year>. </mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X. X.</given-names> <surname>Zhu</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Ramanan</surname></string-name></person-group>, &#x201C;<article-title>Face detection, pose estimation, and landmark localization in the wild</article-title>,&#x201D; in <conf-name>Proc. CVPR</conf-name>, <publisher-loc>Providence, USA</publisher-loc>, pp. <fpage>2879</fpage>&#x2013;<lpage>2886</lpage>, <year>2012</year>. </mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>C. C.</given-names> <surname>Loy</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Tang</surname></string-name></person-group>, &#x201C;<article-title>Wider face: A face detection benchmark</article-title>,&#x201D; in <conf-name>Proc. CVPR</conf-name>, <publisher-loc>Lasvegas, USA</publisher-loc>, pp. <fpage>5525</fpage>&#x2013;<lpage>5533</lpage>, <year>2016</year>. </mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Dai</surname></string-name>, <string-name><given-names>X. R.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>P. S.</given-names> <surname>Chang</surname></string-name> and <string-name><given-names>X. Z.</given-names> <surname>He</surname></string-name></person-group>, &#x201C;<article-title>RSOD: Real-time small object detection algorithm in UAV-based traffic monitoring</article-title>,&#x201D; <source>Applied Intelligence</source>, vol. <volume>1</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Rokach</surname></string-name></person-group>, &#x201C;<article-title>Ensemble-based classifier</article-title>,&#x201D; <source>Artificial Intelligence Review</source>, vol. <volume>33</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>39</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Razavian</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Azizpour</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Sullivan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Carlsson</surname></string-name></person-group>, &#x201C;<article-title>CNN features off-the-shelf: An astounding baseline for recognition</article-title>,&#x201D; in <conf-name>Proc. CVPR</conf-name>, <publisher-loc>Columbus, Ohio, USA</publisher-loc>, pp. <fpage>806</fpage>&#x2013;<lpage>813</lpage>, <year>2014</year>. </mixed-citation></ref>
</ref-list>
</back>
</article>
