<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">IASC</journal-id>
<journal-id journal-id-type="nlm-ta">IASC</journal-id>
<journal-id journal-id-type="publisher-id">IASC</journal-id>
<journal-title-group>
<journal-title>Intelligent Automation &#x0026; Soft Computing</journal-title>
</journal-title-group>
<issn pub-type="epub">2326-005X</issn>
<issn pub-type="ppub">1079-8587</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">27659</article-id>
<article-id pub-id-type="doi">10.32604/iasc.2023.027659</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Low-Cost Real-Time Automated Optical Inspection Using Deep Learning and Attention Map</article-title><alt-title alt-title-type="left-running-head">Low-Cost Real-Time Automated Optical Inspection Using Deep Learning and Attention Map</alt-title><alt-title alt-title-type="right-running-head">Low-Cost Real-Time Automated Optical Inspection Using Deep Learning and Attention Map</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Shih</surname><given-names>Yu</given-names></name>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kuo</surname><given-names>Chien-Chih</given-names></name>
</contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Lee</surname><given-names>Ching-Hung</given-names></name><email>chleenctu@nctu.edu.tw</email>
</contrib><aff><institution>Department of Electrical and Computer Engineering, National Yang Ming Chiao Tung University</institution>, <addr-line>Hsinchu, 300</addr-line>, <country>Taiwan</country></aff>
</contrib-group><author-notes><corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Ching-Hung Lee. Email: <email>chleenctu@nctu.edu.tw</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-06-20"><day>20</day>
<month>06</month>
<year>2022</year></pub-date>
<volume>35</volume>
<issue>2</issue>
<fpage>2087</fpage>
<lpage>2099</lpage>
<history>
<date date-type="received"><day>23</day><month>1</month><year>2022</year></date>
<date date-type="accepted"><day>19</day><month>4</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Shih, Kuo and Lee</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Shih, Kuo and Lee</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_IASC_27659.pdf"></self-uri>
<abstract>
<p>The recent trends in Industry 4.0 and Internet of Things have encouraged many factory managers to improve inspection processes to achieve automation and high detection rates. However, the corresponding cost results of sample tests are still used for quality control. A low-cost automated optical inspection system that can be integrated with production lines to fully inspect products without adjustments is introduced herein. The corresponding mechanism design enables each product to maintain a fixed position and orientation during inspection to accelerate the inspection process. The proposed system combines image recognition and deep learning to measure the dimensions of the thread and identify its defects within 20 s, which is lower than the production-line productivity per 30 s. In addition, the system is designed to be used for monitoring production lines and equipment status. The dimensional tolerance of the proposed system reaches 0.012 mm, and a 100% accuracy is achieved in terms of the defect resolution. In addition, an attention-based visualization approach is utilized to verify the rationale for the use of the convolutional neural network model and identify the location of thread defects.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Automated optical inspection</kwd>
<kwd>deep learning</kwd>
<kwd>real-time inspection</kwd>
<kwd>attention</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>As the basis of Industry 4.0, automated inspection is widely performed in various manufacturing applications to ensure consistent product quality [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]. However, manual and sampling inspections remain prevalent in quality control for reducing manufacturing costs; additionally, monitoring systems to reduce losses caused by machine tool abnormalities are insufficient. In previous studies, the importance of system integration and costs are not considered in the system design. Consequently, a low-cost automated optical inspection (AOI) system that integrates a production line to inspect products completely without adjustment is required. In addition, the results of automated full inspection or product monitoring can be used to directly diagnose the tool wear status of machine tools directly [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. This provides a low-cost monitoring function for ensuring product quality without installing expensive monitoring equipment. This study focuses on the development of an AOI system using deep learning for the real-time inspection of screws. An automated mechanism was designed, and inspection image recognition was used to identify the size and thread defects in screws (as shown in <xref ref-type="fig" rid="fig-1">Figs. 1</xref> and <xref ref-type="fig" rid="fig-2">2</xref>), thereby ensuring consistent product quality.</p>

<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Technical drawing of a screw</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-1.png"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Detection results of (a) normal (qualified) and (b) defective thread</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-2.png"/>
</fig> 
<p>In automation, the mode of transportation affects the speed and efficiency of inspection. Currently, production line transportation is typically performed using robot manipulators [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. Robot manipulators can improve the adjustability of the process; however, calibration for different tasks is time consuming, particularly in collaborative robotic systems. Another typical transportation method involves automatic guided vehicles (AGVs) [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>]. AGVs with suitable path-planning algorithms used in intelligent workshops render material transportation more efficient. However, AGVs are more suitable for transporting bulk or large materials. Moreover, companies with low capital cannot afford the high costs of robots and land. Because the conveyor is suitable for transporting unconfined materials [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>] and does not require calibration, it is typically used in workshops. To establish an inspection system rapidly, a conveyor was adopted in this study. In addition, image histogram equalization and Canny edge detection are typically used to improve the detection accuracy of image recognition systems for determining the dimensions of an object [<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-22">22</xref>]. Recently, detection consistency has been improved using deep learning methods, particularly convolutional neural networks (CNNs) [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. Among CNNs, VGG16 is trained using one million images, which contain 1000 categories that encompass almost all objects in daily life [<xref ref-type="bibr" rid="ref-29">29</xref>]. The complex structure and significant amount of data of VGG16 endow it with high feature extraction ability. In this study, VGG16 with transfer learning was used to establish the proposed defect detection system. To verify the detection results, two visualization methods with attention mapping, gradient-weighted class activation mapping (Grad-CAM) [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>], and Grad-CAM&#x002B;&#x002B; [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>] were used to generate attention maps. Unlike previous visualization methods [<xref ref-type="bibr" rid="ref-34">34</xref>&#x2013;<xref ref-type="bibr" rid="ref-36">36</xref>], Grad-CAM does not require model structure changes to provide effective explanations. The areas of interest of the model are denoted in the attention maps with weights. Grad-CAM&#x002B;&#x002B; improves the weight <inline-formula id="ieqn-1">
<mml:math id="mml-ieqn-1"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:math>
</inline-formula>, which considers the pixel location. In this study, actual differences caused by pixel location abnormalities are compared. Attention maps are used to denote the locations of defects that assist in the reverification process.</p>
<p>System integration is a complex process for achieving intelligent manufacturing, particularly for workshops with low capital. To reduce production costs, the equipment used often spans multiple generations and is sourced from different suppliers, rendering system integration difficult. Therefore, an add-on quality control system is introduced to provide the maximum production benefits. The design of the mechanism allows the AOI system to be connected to the production line without additional calibration. In addition, the transfer learning method for the defect detection model can reduce the number of computer calculations required. The full evaluation of the dimension and the defect detection results can be recorded via real-time full inspection; as such, condition monitoring without requiring numerous sensors in the equipment is achieved. Finally, the results of the attention map verify the rationale for using the CNN model and denoting the location of thread defects, which improve the credibility of the model.</p>
<p>The remainder of this paper is organized as follows: Section 2 introduces the system mechanism and hardware specifications. The established deep learning defect detection method is introduced in Section 3. Section 4 presents the corresponding experimental results and discussion. Finally, the conclusions are presented in Section 5.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Mechanism Design</title>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the mechanism of the automated real-time inspection system. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the proposed system includes an input conveyor, position grooves, an inspection area, and a classification region. To achieve real-time inspection, two conveyors are used to integrate the proposed system with the machining process. Screw positioning, image capture, and recognition are achieved for each screw during movement. The proposed system can be placed directly at the end of the production line without adjustment.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Mechanism of automated real-time inspection</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-3.png"/>
</fig>
<p>The feeding conveyor transports the screws from the machining process to the positioning mechanism. The feeding conveyor belt, which is flat and high speed, is used to move the workpieces. To ensure stability during inspection, a double-sided toothed belt is used in the detection conveyor to fix the screw in place, thereby preventing slippage between the belt and pulley. In addition, a pad is placed in the middle of the detection conveyor to roll screws over this area; thus, the upper camera can capture multiple images with different thread angles to ensure the integrity of detection. The mechanism for rolling the screw as it passes through the upper camera is illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<p>In this study, a mechanism that does not require additional power is designed to fix the direction and position of the screw. The mechanism comprises a landslide and an aluminum block. The shape of the landslide is similar to that of a funnel, which fixes the screw to the positioning block in a particular direction. Subsequently, the weight of the aluminum block enables the screw to be turned over and placed securely on the detection conveyor when it enters the aluminum block. After the screw departs from the aluminum block, it automatically returns to its original position. The positioning process can be completed without additional power sources or sensors.</p>
<p>The light source directly affects the image quality. In this study, four LED strips were arranged in a square to achieve an effect similar to that of a ring light. In the classification area, four servomotors were used to control the four barriers. The classification area was classified into one area for qualified screws and three areas for unqualified screws, namely those with a large diameter, a small diameter, or thread defects. After an object is inspected, a barrier opens based on its classification, and the other barriers close, forming a slope that enables the screws to enter the appropriate area. The system is integrated with Arduino Uno programmed using Python and C#. The Arduino Uno is matched with L298N to operate the motor at the required speed. To facilitate operation, the camera, motor, and light source are integrated with the C# user interface through the Arduino Uno. In addition, the interface communicates with Python to update the dimension measurements and defect detection results. The proposed system combines image recognition and deep learning to measure the dimensions of the thread and identify its defects within 20 s, which is lower than the production-line productivity per 30 s. Thus, the AOI system can achieve real-time quality control.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Mechanism of rolling the screw as it passes through upper camera</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-4.png"/>
</fig>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Convolutional neural network structure</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-5.png"/>
</fig>
</sec>
<sec id="s3">
<label>3</label>
<title>Defect Detection Using Deep Learning with Attention Map</title>
<p>The proposed defect-detection method using a CNN with an attention map is introduced in this section. Grad-CAM&#x002B;&#x002B; was adopted to create the defect region for monitoring and reverification. First, we used discontinuous edge images to determine the dimensions of the recognized objects. After image preprocessing, the Sobel operator and Canny edge detection were used to identify boundaries by adjusting the threshold.</p>
<sec id="s3_1">
<label>3.1</label>
<title>CNN</title>
<p>Neural networks can mitigate fluctuations in inspection results caused by the manual inspection of abundant data; meanwhile, extensive calculations, which are performed via mathematical or computational models that mimic the structure of biological neural networks, are performed in machine learning and cognitive science to approximate neurological functions. Neural networks are adaptive systems, and the most typically used type of deep learning method in image processing is CNNs [<xref ref-type="bibr" rid="ref-23">23</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>]. CNNs are advantageous owing to their capability to automatically extract features from images. CNNs comprise three layers, i.e., convolutional, pooling, and fully connected layers, as illustrated in the architecture shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Convolutional Layer</title>
<p>The convolutional layer contains a filter matrix for calculating the output neurons via local input weights and the connected region. In a grayscale image, the convolutional operation is expressed as follows:</p>
<p><disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mo>&#x2297;</mml:mo><mml:mi>K</mml:mi><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:msubsup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mrow><mml:mrow><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-2">
<mml:math id="mml-ieqn-2"><mml:mi>I</mml:mi><mml:mo>&#x2297;</mml:mo><mml:mi>K</mml:mi></mml:math>
</inline-formula> denotes the convolutional operations of the images and kernels; <italic>m</italic> and <italic>n</italic> denote the pixels in the <italic>m-</italic>th row and <italic>n</italic>th column, respectively. Subsequently, after convolution, a nonlinear activation function such as a sigmoid or rectified linear unit is used to determine the output (feature map), as follows:</p>
<p><disp-formula id="eqn-2"><label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:msubsup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mrow><mml:msubsup><mml:mi>K</mml:mi><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mrow><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Pooling Layer</title>
<p>The pooling layer is used to reduce the size of the input image. In the pooling process, the convolutional feature matrix is partitioned into regions, and the maximum or average values of each region are obtained. The pooling operation is expressed as</p>
<p><disp-formula id="eqn-3"><label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>Extracting representative features from the pooling layer significantly reduces the number of parameters.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Fully Connected Layer</title>
<p>The fully connected layer calculates the classification score and determines the final category. The operation of the <italic>j-</italic>th neuron in the fully connected layer <italic>L</italic> is expressed as</p>
<p><disp-formula id="eqn-4"><label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mi>L</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>L</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>j</mml:mi><mml:mi>L</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-3">
<mml:math id="mml-ieqn-3"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math>
</inline-formula> denotes the input of the <italic>j-</italic>th neuron, <inline-formula id="ieqn-4">
<mml:math id="mml-ieqn-4"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="normal">w</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula> the weight of the input <inline-formula id="ieqn-7">
<mml:math id="mml-ieqn-7"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="normal">x</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">L</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math>
</inline-formula>, <inline-formula id="ieqn-9">
<mml:math id="mml-ieqn-9"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">L</mml:mi></mml:mrow></mml:mrow></mml:msubsup></mml:math>
</inline-formula> the bias, <italic>f</italic> the activation function, and <inline-formula id="ieqn-10">
<mml:math id="mml-ieqn-10"><mml:msubsup><mml:mi>y</mml:mi><mml:mi>j</mml:mi><mml:mi>L</mml:mi></mml:msubsup></mml:math>
</inline-formula> the corresponding output.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Attention-based CNN</title>
<p>In numerous deep-learning models, multilayer networks are used to automatically select features to achieve high accuracy. In most CNNs, global average pooling (GAP) is used in class activation mapping (CAM) to replace the fully connected layer, enabling the model to support inputs of any size and retain abundant information after multiple convolutions and pooling. Recently, the global average of the gradient is used in Grad-CAM to calculate the weights from the feature maps to overcome the limitations of CAM [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. In Grad-CAM, the weight of the <italic>k</italic>th feature map of category <italic>c</italic> is defined as follows:</p>
<p><disp-formula id="eqn-5"><label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac></mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <italic>Z</italic> denotes the number of pixels in the feature map, <inline-formula id="ieqn-11">
<mml:math id="mml-ieqn-11"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:math>
</inline-formula> the score of the corresponding category <italic>c</italic>, and <inline-formula id="ieqn-12">
<mml:math id="mml-ieqn-12"><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:math>
</inline-formula> the pixel value at position (<italic>i</italic>, <italic>j</italic>) in the <italic>k</italic>th feature map. After obtaining the category weights of all feature maps, the weighted sum is calculated to obtain the attention map (or heat map). The attention map can be represented as</p>
<p><disp-formula id="eqn-6"><label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:mrow><mml:mrow><mml:msup><mml:mi>A</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula></p>
<p>The attention map obtained using Grad-CAM increases the transparency of the CNN model. Although Grad-CAM provides the visualization of CNN for classification, it cannot accurately locate objects when the input images contain multiple objects in the same class [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. Therefore, to identify object locations more accurately, the weight formulation connected to the feature map should be modified. The weights <inline-formula id="ieqn-13">
<mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:math>
</inline-formula> independent of position (<italic>i</italic>, <italic>j</italic>) in Grad-CAM can be represented as shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>. In Grad-CAM, every sample is composed of multiple input image categories; therefore, its formulation can be used to identify the target location without the position (<italic>i</italic>, <italic>j</italic>). However, the pixel position is critical for an input image containing multiple objects in the same class. In Grad-CAM&#x002B;&#x002B;, the pixel locations are considered in the weights <inline-formula id="ieqn-14">
<mml:math id="mml-ieqn-14"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:math>
</inline-formula> [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>], and the calculation formula is</p>
<p><disp-formula id="eqn-7"><label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mspace width="0.056em" /></mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mspace width="0.056em" /></mml:mrow><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-15">
<mml:math id="mml-ieqn-15"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:math>
</inline-formula> denotes the score of class <italic>c</italic>, and <inline-formula id="ieqn-16">
<mml:math id="mml-ieqn-16"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> denotes the gradient weights. <inline-formula id="ieqn-17">
<mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> can be obtained by partially differentiating <inline-formula id="ieqn-18">
<mml:math id="mml-ieqn-18"><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:math>
</inline-formula> and <inline-formula id="ieqn-19">
<mml:math id="mml-ieqn-19"><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:math>
</inline-formula> twice, and <inline-formula id="ieqn-20">
<mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> can be represented as</p>
<p><disp-formula id="eqn-8"><label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>a</mml:mi></mml:msub><mml:mrow><mml:mspace width="0.056em" /></mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>b</mml:mi></mml:msub><mml:mrow><mml:mspace width="0.056em" /></mml:mrow><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mn>3</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msup><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msubsup><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where (<italic>i</italic>, <italic>j</italic>) and (<italic>a</italic>, <italic>b</italic>) denote the same iterators in the <italic>k</italic>th feature map, and <inline-formula id="ieqn-21">
<mml:math id="mml-ieqn-21"><mml:mrow><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula> is the solution for <inline-formula id="ieqn-22">
<mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> in Grad-CAM&#x002B;&#x002B;. When <inline-formula id="ieqn-23">
<mml:math id="mml-ieqn-23"><mml:msubsup><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math>
</inline-formula>, Grad-CAM&#x002B;&#x002B; reduces to Grad-CAM, i.e., Grad-CAM&#x002B;&#x002B; provides a general formulation for obtaining the attention map of the input image. Grad-CAM and Grad-CAM&#x002B;&#x002B; exhibit the same structure despite their various differences, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Differences between Grad-CAM and Grad-CAM&#x002B;&#x002B;</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-6.png"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Defect Detection System</title>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows the system flowchart of the proposed AOI defect detection. The screw is sent from the machine tool to the positioning mechanism by the feeding conveyor (shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>) and then placed on the detection conveyor. When the screw enters the classification area, the corresponding images from three perspectives are captured, and then measurement and defect detection are performed. The computer completes the image recognition and then sends the results to the classification area, where the screw is sorted. After the screw is sorted, the results are shown in the user interface (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>) and the webpage for monitoring (<xref ref-type="fig" rid="fig-9">Fig. 9</xref>). The detection results and thread defects are updated at the interface after each inspection, and the images are updated simultaneously. Dimension measurements can assist in evaluating the tool-wear condition of the machine tool. The functions of real-time monitoring, alarm, and suspension systems are provided. The user interface shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref> presents the dimension measurements and defect detection results. In the defect detection results, the output value of a qualified thread is 1, and the output value of a failed thread is 0 (as shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>). The red lines indicate the range of acceptable tolerance. Both the interface and webpage can assist in determining whether the machine is operating abnormally.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>System flowchart for inspection detection</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-7.png"/>
</fig>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>User interface of system</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-8.png"/>
</fig>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Webpage for monitoring machine tools</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-9.png"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results of Defect Detection</title>
<p>Herein, we consider a scenario in which the acquisition and labeling of numerous images is challenging. VGG16, which exhibits high feature extraction capabilities, was utilized and fine-tuned via transfer learning [<xref ref-type="bibr" rid="ref-29">29</xref>], and the corresponding structure is shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. Because the number of defect samples is typically small, VGG16 with parameter transfer learning for defect detection was employed in this study. VGG16 comprises multiple convolutional layers, which endows it with better feature-extraction capabilities compared with ordinary neural networks. VGG16 uses millions of images to classify thousands of categories through 13 convolutional layers and emphasizes the use of numerous 3 &#x00D7; 3 filters in its convolutional layers. When larger filters are used instead of smaller filters, the receptive field is improved, which increases the amount of information obtained. In this study, the output layer of VGG16 was rewritten to classify the two categories regardless of whether a thread contained a defect.</p>

<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>VGG16 structure [<xref ref-type="bibr" rid="ref-29">29</xref>]</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-10.png"/>
</fig>
<p>In this study, an insufficient image was augmented via the translation, rotation, and flipping of the screw. The model was trained for 1000 generations, and the training data included 242 and 228 qualified and defective components. Furthermore, the model achieved 100% accuracy. The laptop computer used in the proposed system comprised an Intel Core i5-5200U processor, a GeForce 930M graphics processing unit, and 8 GB of random-access memory. Test data were obtained from the actual machining results, including those of 20 qualified and 20 defective components. To avoid data imbalance and insufficient data in the training model, only 40 screws were used for the test data. The confusion matrix presented in <xref ref-type="fig" rid="fig-11">Fig. 11</xref> indicates that 100% accuracy was achieved after the actual testing of the system. The corresponding accuracy, precision, and recall rates were all 100%. These results demonstrate the effectiveness of the proposed approach.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Confusion matrix of inspection results based on test data</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-11.png"/>
</fig>
<p>As indicated above, the established model achieved high accuracy despite the low amount of training data used. To ensure that the model accurately identifies defects and facilitates in reverification, Grad-CAM methods were utilized to visualize the defect. The corresponding weight attention maps for defect detection calculated using Grad-CAM and Grad-CAM&#x002B;&#x002B; are presented in <xref ref-type="fig" rid="fig-12">Figs. 12</xref> and <xref ref-type="fig" rid="fig-13">13</xref>, respectively. As shown by the heat map in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>, the attention map prevents the defect region even when the classification accuracy is 100%, i.e., the attention map does not directly indicate defects by selecting the defect area. All test data exhibited the same phenomena. Therefore, Grad-CAM&#x002B;&#x002B; was adopted to improve the localization of the defects. The corresponding attention map obtained using Grad-CAM&#x002B;&#x002B; is shown in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>, where the defect region is denoted. Thus, the judgment of the model is reasonable. Based on this test, the effects of the weights <inline-formula id="ieqn-24">
<mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>w</mml:mi><mml:mi>k</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:math>
</inline-formula> considering the pixel location can be observed. Grad-CAM&#x002B;&#x002B; not only confirms the rationality of the model, but also provides indicators for the entire thread.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Defective thread after performing Grad-CAM</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-12.png"/>
</fig>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Defective thread after performing Grad-CAM&#x002B;&#x002B;</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="IASC_27659-fig-13.png"/>
</fig>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>The proposed system was operated at minimal cost and integrated into a production line without the adjustment of existing processes. Using a five-megapixel camera, a dimensional tolerance of 0.012 mm was obtained, and a defect detection accuracy of 100% was achieved using a limited number of samples. Furthermore, the user interface and webpage can be used to monitor the production lines and equipment status. Additionally, the developed model was visualized using Grad-CAM and Grad-CAM&#x002B;&#x002B;, and the results of the two methods were compared. Grad-CAM&#x002B;&#x002B; performed more effectively, provided a reasonable explanation for each classification result, and accurately located the thread defects. The proposed system resulted in a more complete process and increased consumer confidence as it was able to perform quality control within 20 s.</p>
</sec>
</body>
<back>
<ack>
<p>This study was supported partially by the Ministry of Science and Technology, Taiwan, under contracts MOST-110-2634-F-009-024, 109-2218-E-150-002, and 109-2218-E-005-015.</p>
</ack><fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Hsu</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Shen</surname></string-name></person-group>, &#x201C;<article-title>The design and implementation of an embedded real-time automated IC marking inspection system</article-title>,&#x201D; <source>IEEE Transactions on Semiconductor Manufacturing</source>, vol. <volume>32</volume>, no. <issue>1</issue>, pp. <fpage>112</fpage>&#x2013;<lpage>120</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Jeon</surname></string-name>, <string-name><given-names>U.</given-names> <surname>Jung</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kim</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Han</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Vision-inspection-synchronized dual optical coherence tomography for high-resolution real-time multidimensional defect tracking in optical thin film industry</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>190700</fpage>&#x2013;<lpage>190709</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Gao</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>A new ensemble approach based on deep convolutional neural networks for steel surface defect classification</article-title>,&#x201D; in <conf-name>Proc. 51st CIRP Conf. on Manufacturing Systems</conf-name>, <publisher-loc>Stockholm, Sweden, vol. </publisher-loc><volume>72</volume>, pp. <fpage>1069</fpage>&#x2013;<lpage>1072</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Mei</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Hu</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Wen</surname></string-name></person-group>, &#x201C;<article-title>Deep learning based automated inspection of weak microscratches in optical fiber connector end-face</article-title>,&#x201D; <source>IEEE Transactions on Instrumentation and Measurement</source>, vol. <volume>70</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Luo</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Fang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Automated visual defect detection for flat steel surface: A survey</article-title>,&#x201D; <source>IEEE Transactions on Instrumentation and Measurement</source>, vol. <volume>69</volume>, no. <issue>3</issue>, pp. <fpage>626</fpage>&#x2013;<lpage>644</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Imoto</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Nakai</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Ike</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Haruki</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Sato</surname></string-name></person-group>, &#x201C;<article-title>A CNN-based transfer learning method for defect classification in semiconductor manufacturing</article-title>,&#x201D; <source>IEEE Transactions on Semiconductor Manufacturing</source>, vol. <volume>32</volume>, no. <issue>4</issue>, pp. <fpage>455</fpage>&#x2013;<lpage>459</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Pang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>A new image recognition and classification method combining transfer learning algorithm and mobilenet model for welding defects</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>119951</fpage>&#x2013;<lpage>119960</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Ma</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Zheng</surname></string-name></person-group>, &#x201C;<article-title>A unified system residual life prediction method based on selected tribodiagnostic data</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>44087</fpage>&#x2013;<lpage>44096</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Hao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bian</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Gebraeel</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Shi</surname></string-name></person-group>, &#x201C;<article-title>Residual life prediction of multistage manufacturing processes with interaction between tool wear and product quality degradation</article-title>,&#x201D; <source>IEEE Transactions on Automation Science and Engineering</source>, vol. <volume>14</volume>, no. <issue>2</issue>, pp. <fpage>1211</fpage>&#x2013;<lpage>1224</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Qian</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Yao</surname></string-name></person-group>, &#x201C;<article-title>A symbolic regression based residual useful life model for slewing bearings</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>72076</fpage>&#x2013;<lpage>72089</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>An incidental delivery based method for resolving multirobot pairwised transportation problems</article-title>,&#x201D; <source>IEEE Transactions on Intelligent Transportation Systems</source>, vol. <volume>17</volume>, no. <issue>7</issue>, pp. <fpage>1852</fpage>&#x2013;<lpage>1866</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Hichri</surname></string-name>, <string-name><given-names>J. C.</given-names> <surname>Fauroux</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Adouane</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Doroftei</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Mezouar</surname></string-name></person-group>, &#x201C;<article-title>Design of cooperative mobile robots for co-manipulation and transportation tasks</article-title>,&#x201D; <source>Robotics and Computer-Integrated Manufacturing</source>, vol. <volume>57</volume>, no. <issue>3</issue>, pp. <fpage>412</fpage>&#x2013;<lpage>421</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Hawley</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Suleiman</surname></string-name></person-group>, &#x201C;<article-title>Control framework for cooperative object transportation by two humanoid robots</article-title>,&#x201D; <source>Robotics and Autonomous Systems</source>, vol. <volume>115</volume>, no. <issue>5&#x2013;6</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I. G.</given-names> <surname>Plaksina</surname></string-name>, <string-name><given-names>G. I.</given-names> <surname>Chistokhina</surname></string-name> and <string-name><given-names>D. V.</given-names> <surname>Topolskiy</surname></string-name></person-group>, &#x201C;<article-title>Development of a transport robot for automated warehouses</article-title>,&#x201D; in <conf-name>2018 Int. Multi-Conf. on Industrial Engineering and Modern Technologies (FarEastCon)</conf-name>, <publisher-loc>Vladivostok, Russia</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Nishi</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Akiyama</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Higashi</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Kumagai</surname></string-name></person-group>, &#x201C;<article-title>Cell-based local search heuristics for guide path design of automated guided vehicle systems with dynamic multicommodity flow</article-title>,&#x201D; <source>IEEE Transactions on Automation Science and Engineering</source>, vol. <volume>17</volume>, no. <issue>2</issue>, pp. <fpage>966</fpage>&#x2013;<lpage>980</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Rozsa</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Sziranyi</surname></string-name></person-group>, &#x201C;<article-title>Obstacle prediction for automated guided vehicles based on point clouds measured by a tilted LIDAR sensor</article-title>,&#x201D; <source>IEEE Transactions on Intelligent Transportation Systems</source>, vol. <volume>19</volume>, no. <issue>8</issue>, pp. <fpage>2708</fpage>&#x2013;<lpage>2720</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Hou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Song</surname></string-name></person-group>, &#x201C;<article-title>Research on multi-AGVs path planning and coordination mechanism</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>213345</fpage>&#x2013;<lpage>213356</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Lyons</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Bierie</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Marti</surname></string-name></person-group>, &#x201C;<article-title>Belt conveyor training: Changing behaviors, reducing risk, improving the bottom line</article-title>,&#x201D; in <conf-name>2015 IEEE-IAS/PCA Cement Industry Conf. (IAS/PCA CIC)</conf-name>, <publisher-loc>Toronto, ON, Canada</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2015</year>. </mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Qu</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Qiao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Pang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Research on ADCN method for damage detection of mining conveyor belt</article-title>,&#x201D; <source>IEEE Sensors Journal</source>, vol. <volume>21</volume>, no. <issue>6</issue>, pp. <fpage>8662</fpage>&#x2013;<lpage>8669</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>R. C.</given-names> <surname>Gonzalez</surname></string-name> and <string-name><given-names>R. E.</given-names> <surname>Woods</surname></string-name></person-group>, <source>Digital Image Processing</source>, <edition>3rd</edition> ed., <publisher-name>Pearson Education international, Upper Saddle River</publisher-name>, <publisher-loc>NJ</publisher-loc>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Guo</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Contrast enhancement using stratified parametric-oriented histogram equalization</article-title>,&#x201D; <source>IEEE Transactions on Circuits and Systems for Video Technology</source>, vol. <volume>27</volume>, no. <issue>6</issue>, pp. <fpage>1171</fpage>&#x2013;<lpage>1181</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Canny</surname></string-name></person-group>, &#x201C;<article-title>A computational approach to edge detection</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>PAMI-8</volume>, no. <issue>6</issue>, pp. <fpage>679</fpage>&#x2013;<lpage>698</lpage>, <year>1986</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Ashwini</surname></string-name> and <string-name><given-names>S. B.</given-names> <surname>Rudraswamy</surname></string-name></person-group>, &#x201C;<article-title>Automated inspection system for automobile bearing seals</article-title>,&#x201D; <source>Materials Today: Proceedings</source>, vol. <volume>46</volume>, no. <issue>10</issue>, pp. <fpage>4709</fpage>&#x2013;<lpage>4715</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Fotouhi</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Pashmforoush</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bodaghi</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Fotouhi</surname></string-name></person-group>, &#x201C;<article-title>Autonomous damage recognition in visual inspection of laminated composite structures using deep learning</article-title>,&#x201D; <source>Composite Structures</source>, vol. <volume>268</volume>, no. <issue>3</issue>, pp. <fpage>113960</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W. W.</given-names> <surname>Fan</surname></string-name> and <string-name><given-names>C. H.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>Classification of imbalanced data using deep learning with adding noise</article-title>,&#x201D; <source>Journal of Sensors</source>, vol. <volume>2021</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>18</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>He</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Meng</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Yan</surname></string-name></person-group>, &#x201C;<article-title>An end-to-end steel surface defect detection approach via fusing multiple hierarchical features</article-title>,&#x201D; <source>IEEE Transactions on Instrumentation and Measurement</source>, vol. <volume>69</volume>, no. <issue>4</issue>, pp. <fpage>1493</fpage>&#x2013;<lpage>1504</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Shi</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>A layer-wise multi-defect detection system for powder bed monitoring: lighting strategy for imaging, adaptive segmentation and classification</article-title>,&#x201D; <source>Materials &#x0026; Design</source>, vol. <volume>210</volume>, no. <issue>1&#x2013;4</issue>, pp. <fpage>110035</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Lecun</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bottou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Bengio</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Haffner</surname></string-name></person-group>, &#x201C;<article-title>Gradient-based learning applied to document recognition</article-title>,&#x201D; <source>Proceedings of the IEEE</source>, vol. <volume>86</volume>, no. <issue>11</issue>, pp. <fpage>2278</fpage>&#x2013;<lpage>2324</lpage>, <year>1998</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Simonyan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Very deep convolutional networks for large-scale image recognition</article-title>,&#x201D; in <conf-name>2015 Int. Conf. on Learning Representations (ICLR)</conf-name>, <publisher-loc>San Diego, CA, USA</publisher-loc>, <year>2015</year>. </mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R. R.</given-names> <surname>Selvaraju</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Cogswell</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Das</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Vedantam</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Parikh</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Grad-CAM: visual explanations from deep networks via gradient-based localization</article-title>,&#x201D; in <conf-name>2017 IEEE Int. Conf. on Computer Vision (ICCV)</conf-name>, <publisher-loc>Venice, Italy</publisher-loc>, pp. <fpage>618</fpage>&#x2013;<lpage>626</lpage>, <year>2017</year>. </mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. Y.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>C. H.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>Vibration signals analysis by explainable artificial intelligence (XAI) approach: Application on bearing faults diagnosis</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>134246</fpage>&#x2013;<lpage>134256</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Chattopadhay</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sarkar</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Howlader</surname></string-name> and <string-name><given-names>V. N.</given-names> <surname>Balasubramanian</surname></string-name></person-group>, &#x201C;<article-title>Grad-CAM&#x002B;&#x002B;: Generalized gradient-based visual explanations for deep convolutional networks</article-title>,&#x201D; in <conf-name>2018 IEEE Winter Conf. on Applications of Computer Vision (WACV)</conf-name>, <publisher-loc>Lake Tahoe, NV, USA</publisher-loc>, pp. <fpage>839</fpage>&#x2013;<lpage>847</lpage>, <year>2018</year>. </mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. R.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>M. C.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>Robust tool wear monitoring system development by sensors and feature fusion</article-title>,&#x201D; <source>Asian Journal of Control</source>, vol. <volume>6</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>17</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. D.</given-names> <surname>Zeiler</surname></string-name>, <string-name><given-names>G. W.</given-names> <surname>Taylor</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Fergus</surname></string-name></person-group>, &#x201C;<article-title>Adaptive deconvolutional networks for mid and high level feature learning</article-title>,&#x201D; in <conf-name>2011 Int. Conf. on Computer Vision</conf-name>, <publisher-loc>Barcelona, Spain</publisher-loc>, pp. <fpage>2018</fpage>&#x2013;<lpage>2025</lpage>, <year>2011</year>. </mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. H.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Huh</surname></string-name>, <string-name><given-names>B. A.</given-names> <surname>Tama</surname></string-name>, <string-name><given-names>S. Y.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J. H.</given-names> <surname>Jung</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Vision-based fault diagnostics using explainable deep learning with class activation maps</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>129169</fpage>&#x2013;<lpage>129179</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Khosla</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Lapedriza</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Oliva</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Torralba</surname></string-name></person-group>, &#x201C;<article-title>Learning deep features for discriminative localization</article-title>,&#x201D; in <conf-name>2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, pp. <fpage>2921</fpage>&#x2013;<lpage>2929</lpage>, <year>2016</year>. </mixed-citation></ref>
</ref-list>
</back>
</article>