<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id> 
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">19082</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2022.019082</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning-Based Automatic Detection and Evaluation on Concrete Surface Bugholes</article-title>
<alt-title alt-title-type="left-running-head">Deep Learning-Based Automatic Detection and Evaluation on Concrete Surface Bugholes</alt-title>
<alt-title alt-title-type="right-running-head">Deep Learning-Based Automatic Detection and Evaluation on Concrete Surface Bugholes</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Wei</surname>
<given-names>Fujia</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="aff" rid="aff-2">2</xref><email>weifujia13@163.com</email>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western">
<surname>Shen</surname>
<given-names>Liyin</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Xiang</surname>
<given-names>Yuanming</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Zhang</surname>
<given-names>Xingjie</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western">
<surname>Tang</surname>
<given-names>Yu</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western">
<surname>Tan</surname>
<given-names>Qian</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>School of Management Science and Real Estate, Chongqing University</institution>, <addr-line>Chongqing, 400044</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>CMCU Engineering Co., Ltd.</institution>, <addr-line>Chongqing, 400039</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Fujia Wei. Email: <email>weifujia13@163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-03-11">
<day>11</day>
<month>03</month>
<year>2022</year>
</pub-date>
<volume>131</volume>
<issue>2</issue>
<fpage>619</fpage>
<lpage>637</lpage>
<history>
<date date-type="received">
<day>01</day>
<month>9</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>05</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Binesh</copyright-statement>
<copyright-year>2022</copyright-year>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_19082.pdf"></self-uri>
<abstract>
<p>Concrete exterior quality is one of the important metrics in evaluating construction project quality. Among the defects affecting concrete exterior quality, bughole is one of the most common imperfections, thus detecting concrete bughole accurately is significant for improving concrete exterior quality and consequently the quality of the whole project. This paper presents a deep learning-based method for detecting concrete surface bugholes in a more objective and automatic way. The bugholes are identified in concrete surface images by Mask R-CNN. An evaluation metric is developed to indicate the scale of concrete bughole. The proposed approach can detect bugholes in an instance level automatically and output the mask of each bughole, based on which the bughole area ratio is automatically calculated and the quality grade of the concrete surfaces is assessed. For demonstration, a total of 273 raw concrete surface images taken by mobile phone cameras are collected as a dataset. The test results show that the average precision (AP) of bughole masks is 90.8&#x025;.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Defect detection</kwd>
<kwd>engineering</kwd>
<kwd>concrete quality</kwd>
<kwd>deep learning</kwd>
<kwd>instance segmentation</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s_1">
<label>1</label>
<title>Introduction</title>
<p>Concrete is the most widely used material in civil engineering structures. When the surface of a building is mainly composed of concrete, a high-quality exterior becomes an important factor of construction quality. Among the defects affecting the exterior quality, bughole is the most typical one [<xref ref-type="bibr" rid="ref-1">1</xref>]. A recent questionnaire indicates that the exterior of concrete is as important as cost-performance attributes and workability [<xref ref-type="bibr" rid="ref-2">2</xref>]. Moreover, the consequences of bugholes could be serious if the concrete surfaces are to be painted or the damaged area reaches a certain threshold [<xref ref-type="bibr" rid="ref-3">3</xref>, <xref ref-type="bibr" rid="ref-4">4</xref>]. Related investigations have shown that bugholes on concrete surfaces affect the subsequent painting construction because these defects need to be filled before painting, which causes additional workload and cost [<xref ref-type="bibr" rid="ref-5">5</xref>]. Therefore, bugholes should be minimized during the construction process to improve the flatness and aesthetics of the concrete structure. Traditional detection methods rely on manual inspection [<xref ref-type="bibr" rid="ref-6">6</xref>, <xref ref-type="bibr" rid="ref-7">7</xref>], which is considered time-consuming and impractical [<xref ref-type="bibr" rid="ref-8">8</xref>, <xref ref-type="bibr" rid="ref-9">9</xref>]. An improved method of bughole rating recommended by both the Concrete International Board (CIB) and American Concrete Institute (ACI) suggests comparing the concrete surfaces with reference bughole photo samples. The scales of reference bughole photo samples are illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. However, the effectiveness of this method is affected by both the printed scales of reference samples and the subjectivity of human inspectors [<xref ref-type="bibr" rid="ref-10">10</xref>, <xref ref-type="bibr" rid="ref-11">11</xref>]. In addition, one surface may have several types of imperfections, hence the effectiveness of reference samples is limited.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>The scales of reference bughole photo samples</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-1.png"/></fig>
<p>With the development of image processing technology [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>], image processing methods have been used in many areas including concrete bridge inspection [<xref ref-type="bibr" rid="ref-12">12</xref>], classification of radar images [<xref ref-type="bibr" rid="ref-13">13</xref>], and so on. To enable the inspectors to use the reference bughole photo samples more objectively, some other studies have proposed image processing techniques to detect and evaluate the distribution of bugholes on concrete surfaces [<xref ref-type="bibr" rid="ref-10">10</xref>, <xref ref-type="bibr" rid="ref-14">14</xref>]. However, some argue that the accuracy of image processing methods is affected by noise such as illumination, shadows, and combinations of several different surface defects [<xref ref-type="bibr" rid="ref-15">15</xref>, <xref ref-type="bibr" rid="ref-16">16</xref>]. In recent years, algorithms based on deep learning have achieved excellent progress in the challenge of object detection [<xref ref-type="bibr" rid="ref-17">17</xref>]. Deep learning is a sub-field of machine learning. It uses many levels of non-linear information processing and abstraction for supervised or unsupervised feature learning and representation, classification, and pattern recognition [<xref ref-type="bibr" rid="ref-18">18</xref>]. Compared with traditional machine learning, deep learning aims to automatically extract multi-layer feature representations from data. Its core idea is to use a series of non-linear transformations in a data-driven way to extract features from the original data from low-level to high-level, from specific to abstract, and from general to specific semantics. That is, deep learning has powerful capabilities and flexibility in supporting computer systems to be improved from experience and data. Some researchers have applied deep learning-based models in the construction industry. Related research focuses on the application of deep learning-based object detection algorithms to identify, classify, and locate damages on structural surfaces, such as crack detection [<xref ref-type="bibr" rid="ref-19">19</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>], concrete spalling detection [<xref ref-type="bibr" rid="ref-25">25</xref>, <xref ref-type="bibr" rid="ref-26">26</xref>], and corrosion detection [<xref ref-type="bibr" rid="ref-27">27</xref>]. Among these research, most studies focus on the use of Convolutional Neural Networks (CNNs) to realize the classification and localization of defects. However, insufficient attention has been paid to the evaluation of concrete surfaces [<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<p>In line with the above research backgrounds, this paper proposes a deep learning-based method using Mask R-CNN [<xref ref-type="bibr" rid="ref-29">29</xref>] to detect bugholes in concrete surface images and support decision-making for quality improvement. The proposed approach can recognize bugholes at an instance level and output the pixel of each bughole. Moreover, the area ratio of bugholes is automatically calculated to evaluate concrete surfaces.</p>
</sec>
<sec id="s_2">
<label>2</label>
<title>Methodology</title>
<p>The overall framework of the proposed bughole detection method is given in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The framework is composed of three stages: 1) database (DB) establishment; 2) establishment of network architecture; 3) bughole detection using trained Mask R-CNN. In the first stage, the camera of a mobile phone is used to acquire images from a fair-faced concrete building under different lighting conditions at distances of 0.1&#x2013;1.0 m. In the second stage, the instance segmentation framework Mask R-CNN is modified to build an end-to-end bughole recognition model. In the third stage, the performance of the trained Mask R-CNN model is evaluated by the test set, and the recognition results are compared with the CIB reference scale to evaluate the bughole rating on the concrete surface. The detailed implementations are described in this section.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Overall framework of the proposed method</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-2.png"/></fig>
<sec id="s_2_1">
<label>2.1</label>
<title>Database Establishment</title>
<p>A mobile phone camera captures images in this study. A total of 273 concrete surface images with a resolution of 3,024 &#x00D7; 3,024 pixels are collected. To create a database to train and test the Mask R-CNN-based detection model introduced in this study, the original images with a resolution of 3,024 &#x00D7; 3,024 pixels are cropped to 256 &#x00D7; 256 pixels, and a total of 3,215 images containing bughole are selected to create the datasets. The number of images in the training set is 2,572, and the number of images in the validation set is 643, according to the ratio of the training set:validation set = 4:1 [<xref ref-type="bibr" rid="ref-24">24</xref>]. Image annotation is the core of semantic object image segmentation in computer vision [<xref ref-type="bibr" rid="ref-30">30</xref>]. This study uses the image annotation tool &#x201C;labelme&#x201d; to annotate the labels and masks of objects (bugholes) [<xref ref-type="bibr" rid="ref-31">31</xref>]. Examples of annotated images are shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Examples of annotated images</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-3.png"/></fig>
<p>A good model requires a lot of data to train, but it is time-consuming and costly to get new data. To overcome this obstacle, data augmentation is a way to increase the amount of data by random scale, crop, flip, shift, noise, and rotation of existing data. Related studies have shown that data augmentation can improve the generalization ability and robustness of the model [<xref ref-type="bibr" rid="ref-20">20</xref>]. In this study, a rotation approach is adopted to augment data. 500 randomly selected images are preprocessed by rotation before the training process. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows examples of image modification for data augmentation.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Examples of data augmentation: (a) original image; (b) rotation 90<sup>&#x2218;</sup> clockwise; (c) flip horizontally; (d) blur; (e) shift; (f) color conversion</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-4.png"/></fig>
</sec>
<sec id="s_2_2">
<label>2.2</label>
<title>Establishment of Network Architecture</title>
<p>Mask R-CNN is an instance segmentation algorithm. It is a more elaborate segmentation process for similar objects based on semantic segmentation. In this study, the Mask R-CNN algorithm framework performs the task of detecting and evaluating bugholes on concrete surface images. The Residual Network (ResNet) [<xref ref-type="bibr" rid="ref-32">32</xref>] with 101 layers and the Feature Pyramid Network (FPN) [<xref ref-type="bibr" rid="ref-33">33</xref>] are selected as the feature extraction network. The reason for using ResNet is to enrich feature extraction by increasing network depth while addressing the degradation problem. FPN can increase the resolution and high-level semantic information of the feature map, thereby improving the object detection performance of the network. The Region Proposal Network (RPN) selects the candidate Region of Interest (RoI) according to different scales, lengths, and widths, and then distinguishes and initially locates multiple RoIs generated on the feature map. The classic object detection algorithm Faster R-CNN [<xref ref-type="bibr" rid="ref-34">34</xref>] classifies individual bugholes and locates each of them by a bounding box. The classical semantic segmentation algorithm fully convolutional network (FCN) [<xref ref-type="bibr" rid="ref-35">35</xref>] generates the corresponding mask branch, which can distinguish each bughole at the instance level. The overall network architecture of the Mask R-CNN framework is shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The specifications of Mask R-CNN without RPN are shown in <xref ref-type="table" rid="table-1">Table 1</xref>. The detailed specification of Conv and identity block with depth (64/64/256) are shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>The architecture of Mask R-CNN</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-5.png"/></fig>
<table-wrap id="table-1"><label>Table 1</label><caption><title>The specification of Mask R-CNN without RPN </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Layer</td>
<td align="left">Type</td>
<td align="left">Depth</td>
<td align="left">Filter size</td>
<td align="left">Stride</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">C1</td>
<td align="left">Zeros padding</td>
<td align="left"></td>
<td align="left"></td>
<td align="left">3</td>
</tr>
<tr>
<td align="left">C1</td>
<td align="left">Conv+BN+ReLU</td>
<td align="left">64</td>
<td align="left">7 &#x00D7; 7</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C1</td>
<td align="left">Max pooling</td>
<td align="left">64</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C2</td>
<td align="left">Conv Block</td>
<td align="left">64/64/256</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C2</td>
<td align="left">Identity Block(&#x00D7;2)</td>
<td align="left">64/64/256</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left">C3</td>
<td align="left">Conv Block</td>
<td align="left">128/128/512</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C3</td>
<td align="left">Identity Block(&#x00D7;3)</td>
<td align="left">128/128/512</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left">C4</td>
<td align="left">Conv Block</td>
<td align="left">256/256/1024</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C4</td>
<td align="left">Identity Block(&#x00D7;22)</td>
<td align="left">256/256/1024</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left">C5</td>
<td align="left">Conv Block</td>
<td align="left">512/512/2048</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left">C5</td>
<td align="left">Identity Block(&#x00D7;2)</td>
<td align="left">512/512/2048</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
</tr>
<tr>
<td align="left">FPN</td>
<td align="left">Conv+Up/Downsample (C1&#x223C; C5)</td>
<td align="left">256</td>
<td align="left">1 &#x00D7; 1/3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left">RoI align</td>
<td align="left"></td>
<td align="left">256</td>
<td align="left">7 &#x00D7; 7/14 &#x00D7; 14</td>
<td align="left"></td>
</tr>
<tr>
<td align="left">Classifier and Bbox head</td>
<td align="left">FC+ReLU</td>
<td align="left">1024</td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"></td>
<td align="left">FC+ReLU</td>
<td align="left">1024</td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Softmax/Regressor</td>
<td align="left">NC/4 &#x00D7; NC</td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">Mask head</td>
<td align="left">Conv+BN+ReLU(&#x00D7;4)</td>
<td align="left">256</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Deconv</td>
<td align="left">256</td>
<td align="left">2 &#x00D7; 2</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Sigmoid</td>
<td align="left">NC</td>
<td align="left">1 &#x00D7; 1</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="tfn1_1">
<p>Note: NC: Number of classes.</p></fn></table-wrap-foot></table-wrap>
<table-wrap id="table-2"><label>Table 2</label><caption><title>The specification of Conv and identity block</title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Block</td>
<td align="left">Type</td>
<td align="left">Depth</td>
<td align="left">Filter size</td>
<td align="left">Stride</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Identity Block</td>
<td align="left">Conv+BN+ReLU</td>
<td align="left">64</td>
<td align="left">1 &#x00D7; 1</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Conv+BN+ReLU</td>
<td align="left">64</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Conv+BN</td>
<td align="left">256</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">1</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">(Add shortcut+ ReLU)</td>
<td align="left">&#x2013;</td>
<td align="left">&#x2013;</td>
<td align="left">&#x2013;</td>
</tr>
<tr>
<td align="left">Conv Block</td>
<td align="left">Conv+BN+ReLU</td>
<td align="left">64</td>
<td align="left">1 &#x00D7; 1</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Conv+BN+ReLU</td>
<td align="left">64</td>
<td align="left">3 &#x00D7; 3</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">Conv+BN</td>
<td align="left">256</td>
<td align="left">1 &#x00D7; 1</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">(Conv+BN)</td>
<td align="left">256</td>
<td align="left">1 &#x00D7; 1</td>
<td align="left">2</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">(Add shortcut + ReLU)</td>
<td align="left">&#x2013;</td>
<td align="left">&#x2013;</td>
<td align="left">&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Since the added Mask branch needs to extract a finer spatial layout of the object, thus it exposes the pixel deviation problem of RoI Pooling in the Faster R-CNN algorithm. To solve the problem, the corresponding RoI alignment strategy is proposed. RoI alignment cancels the quantization operation, calculates four regular sampling points in each bin, calculates the values of these four positions by bilinear interpolation, and then performs the maximum pooling operation so that the pixel mask generated by FCN can retain accurate spatial location, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. The solid-line grid represents the feature map. The orange blocks represent the RoI. RoI alignment operation does not perform quantization on any coordinates involved in the RoI, bins, or sampling points, thus avoiding the misalignment between the RoI and the extracted features caused by quantization [<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Operation of RoI alignment</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-6.png"/></fig>
<p>In the RoI regression process, to make the predicted object window closer to the ground truth box, bounding-box regression is often used for fine-tuning. For windows, a four-dimensional vector (x, y, w, h) is generally used to represent the center point coordinates, width, and height of the window. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the blue dashed window P represents the original proposal, and the red window G represents the ground truth of the object. The goal of RoI regression is to find a relationship that maps the input original window P to a regression window <inline-formula id="ieqn-1"></inline-formula> closer to the ground truth window G.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>Computation graph of bounding box regression</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-7.png"/></fig>
<p>That is: given <italic>P</italic> = (<italic>P<sub>x</sub></italic>, <italic>P<sub>y</sub></italic>, <italic>P<sub>w</sub></italic>, <italic>P<sub>h</sub></italic>,), find a mapping <italic>f</italic> such that: <inline-formula id="ieqn-2"></inline-formula>.</p>
<p>The transformation of bounding-box regression is as follows: 1) Translate the original proposal window <inline-formula id="ieqn-3"></inline-formula>, where <inline-formula id="ieqn-4"></inline-formula> and <inline-formula id="ieqn-5"></inline-formula>, then <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-2">(2)</xref> are obtained:</p>
<disp-formula id="eqn-1">
<label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>G</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math>
</disp-formula>
<disp-formula id="eqn-2">
<label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>G</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math>
</disp-formula>
<p>2) Scale the original proposal window (<italic>S<sub>w</sub></italic>, <italic>S<sub>h</sub></italic>), where <italic>S<sub>w</sub></italic> = <italic>p<sub>w</sub>d<sub>w</sub></italic>(<italic>p</italic>) and <italic>S<sub>h</sub></italic> = <italic>p<sub>h</sub>d<sub>h</sub></italic>(<italic>p</italic>), then <xref ref-type="disp-formula" rid="eqn-3">Eqs. (3)</xref> and <xref ref-type="disp-formula" rid="eqn-4">(4)</xref> are obtained:</p>
<disp-formula id="eqn-3">
<label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>G</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</disp-formula>
<disp-formula id="eqn-4">
<label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>G</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:math>
</disp-formula>
<p>From the above four equations, the translation and scaling required by the predicted proposal are (<italic>d<sub>x</sub></italic>(<italic>p</italic>), <italic>d<sub>y</sub></italic>(<italic>p</italic>), <italic>d<sub>w</sub></italic>(<italic>p</italic>), <italic>d<sub>h</sub></italic>(<italic>p</italic>)), where (<italic>d<sub>x</sub></italic>(<italic>p</italic>), <italic>d<sub>y</sub></italic>(<italic>p</italic>), <italic>d<sub>w</sub></italic>(<italic>p</italic>), <italic>d<sub>h</sub></italic>(<italic>p</italic>)) should be equal to (<italic>t<sub>x</sub></italic>, <italic>t<sub>y</sub></italic>, <italic>t<sub>w</sub></italic>, <italic>t<sub>h</sub></italic>) which is translation and scaling required between ground truth window G and original proposal window P, as shown in <xref ref-type="disp-formula" rid="eqn-5">Eqs. (5)</xref> to <xref ref-type="disp-formula" rid="eqn-8">(8)</xref>:</p>
<disp-formula id="eqn-5">
<label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula>
<disp-formula id="eqn-6">
<label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula>
<disp-formula id="eqn-7">
<label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula>
<disp-formula id="eqn-8">
<label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>G</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math>
</disp-formula>
<p>Therefore, the objective function can be expressed as <inline-formula id="ieqn-6"></inline-formula>, where <inline-formula id="ieqn-7"></inline-formula> is the eigenvector of the input proposal; <italic>w</italic><sub>&#x002A;</sub> <sup><italic>T</italic></sup> is the parameter to be learned; * represents x, y, w, and h, the objective function corresponding to each transformation; <italic>d</italic><sub>&#x002A;</sub>(<italic>p</italic>) represents the coordinates of the prediction window. To minimize the deviation between the prediction window and ground truth window, the loss function is defined in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:</p>
<disp-formula id="eqn-9">
<label>(9)</label>
</disp-formula>
</sec>
<sec id="s_2_3">
<label>2.3</label>
<title>Mask R-CNN Model Training</title>
<p>The Mask R-CNN in this study is trained using a joint training strategy. A total of 100 epochs are trained, in which the entire network is trained with 40 epochs, the feature extraction network module uses 40 epochs for training, and 20 epochs are used to fine-tune the network heads. After the network training is completed, the test set and other original images are used to assess the detection performance of the trained Mask R-CNN. The learning curves of training and validation processes are shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The GPU and CPU training modes are used to train the network. The specific configuration of the system environment is shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>The learning curves of training and validation processes</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-8.png"/></fig>
 
<table-wrap id="table-3"><label>Table 3</label><caption><title>The specific configuration of the system environment </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Configuration</td>
<td align="left">Description</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">RAM</td>
<td align="left">DDR4 16 GB (8G&#x00D7;2)</td>
</tr>
<tr>
<td align="left">CPU</td>
<td align="left">Intel(R) Core i7&#x2013;7700K CPU @4.5 GHz</td>
</tr>
<tr>
<td align="left">GPU</td>
<td align="left">MSI Geforce RTX 2080</td>
</tr>
<tr>
<td align="left">Operating system</td>
<td align="left">Ubuntu</td>
</tr>
<tr>
<td align="left">DL Framework</td>
<td align="left">Keras 2.2.4 and tensorFlow 1.12</td>
</tr>
<tr>
<td align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Before starting training, some parameters need to be specified according to the characteristics of the network architecture and training images, such as batch size, learning rate, momentum, and weight decay. The parameters are specified in <xref ref-type="table" rid="table-4">Table 4</xref>. The training time in GPU mode is much shorter than that in CPU mode. It has been estimated that the total training duration in multi-GPU mode is about 0.67 h, while the total training time in CPU mode is more than 12 h. For the same original image with a resolution of 3024 &#x00D7; 3024 pixels, the detection time in GPU mode is 3 s, and the detection time in CPU mode is 60 s.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>The specific parameters </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Parameters</td>
<td align="left">Description</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Batch size</td>
<td align="left">120</td>
</tr>
<tr>
<td align="left">Learning rate</td>
<td align="left">0.001</td>
</tr>
<tr>
<td align="left">Momentum</td>
<td align="left">0.90</td>
</tr>
<tr>
<td align="left">Weight decay</td>
<td align="left">0.0001</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s_3">
<label>3</label>
<title>Experiment</title>
<sec id="s_3_1">
<label>3.1</label>
<title>Evaluation of the Trained Model</title>
<p>The detection performance of the trained Mask R-CNN model is assessed by Pascal VOC&#x2019;s metric. The average precision (AP) of both bounding boxes and masks of the test set is tested in the case where Intersection-over-Union (IoU) is set to 0.5 and 0.75, respectively. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> illustrates the precision-recall curves of bounding boxes and masks.</p>
<fig id="fig-9"><label>Figure 9</label><caption><title>The precision-recall curves (a) bounding-box (b) mask</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-9.png"/></fig>
<p>As can be seen from <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, different IoU settings affect the average precision of detection. The larger the IoU, the lower the average precision. Conversely, the smaller the IoU, the higher the average precision. In this study, when IoU = 0.5, the average precision of bounding boxes and masks are 0.900 and 0.908, respectively. When IoU = 0.75, the average precision of bounding boxes and masks are 0.675 and 0.647, respectively.</p>
</sec>
<sec id="s_3_2">
<label>3.2</label>
<title>Testing with New Images</title>
<p>A total of 43 raw images (3,024 &#x00D7; 3,024 pixel resolutions) are used for evaluating the bughole detection performance of the proposed method. The AP of bounding boxes and masks for these raw images is recorded and the mean AP (mAP) is computed, as shown in <xref ref-type="fig" rid="fig-10">Figs. 10</xref> and <xref ref-type="fig" rid="fig-11">11</xref>.</p>
<fig id="fig-10"><label>Figure 10</label><caption><title>The AP and mAP of bounding boxes</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-10.png"/></fig>
<fig id="fig-11"><label>Figure 11</label><caption><title>The AP and mAP of masks</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-11.png"/></fig>
<p>As shown in <xref ref-type="fig" rid="fig-10">Figs. 10</xref> and <xref ref-type="fig" rid="fig-11">11</xref>, Image 17 has the highest mAP of 97.5&#x025; and 94.0&#x025; for bounding box and mask, respectively. The bounding box of Image 21 has the lowest AP of 76.9&#x025;. The mask of Image 23 has the lowest AP of 61.1&#x025;. The AP of both bounding boxes and masks of these 43 images at IoU = 0.5 and IoU = 0.75 are listed in <xref ref-type="sec" rid="s_6_1">Appendix A</xref>. <xref ref-type="fig" rid="fig-12">Fig. 12</xref> illustrates the results of bughole detection on concrete surface images by the proposed method. The input is a cropped image with a resolution of 256 &#x00D7; 256 pixels, and the output is the bughole detection result. The first number in the label of the output image represents the probability, and the second number represents the pixel. The bughole recognition results of the sample image in <xref ref-type="fig" rid="fig-12">Fig. 12</xref> are listed in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<fig id="fig-12"><label>Figure 12</label><caption><title>The input and output images</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-12.png"/></fig>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Bughole recognition information </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Detected bughole</td>
<td align="left">No. 1</td>
<td align="left">No. 2</td>
<td align="left">No. 3</td>
</tr>
</thead><tbody>
<tr>
<td align="left">Probability</td>
<td align="left">0.95</td>
<td align="left">0.98</td>
<td align="left">1.00</td>
</tr>
<tr>
<td align="left">Pixels</td>
<td align="left">190.81</td>
<td align="left">330.73</td>
<td align="left">148.93</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s_3_3">
<label>3.3</label>
<title>Evaluation of Bugholes on Concrete Surfaces</title>
<p>Inspired by the CIB bughole rating method, by calculating the percentage of bughole area in the concrete surface image, the surface bugholes of concrete are divided into seven levels. With the aid of the image processing tool, the analysis result of the CIB bughole scale is shown in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>.</p>
<fig id="fig-13"><label>Figure 13</label><caption><title>The result of the CIB bughole scale obtained by image processing</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-13.png"/></fig>
<p>The relationship between the area ratio of surface bugholes and the CIB bughole scale is obtained and shown in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>. The regression analysis equation of the area ratio and CIB bughole scale is shown in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:</p>
<disp-formula id="eqn-10">
<label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mi>B</mml:mi><mml:mi>u</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mo>=</mml:mo><mml:mn>1.3436</mml:mn><mml:mtext>In</mml:mtext><mml:msub><mml:mi>A</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn>10.28</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>0.9817</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>A<sub>b</sub></italic> is the percentage of surface bugholes in the image to be rated and <italic>R</italic> represents the regression coefficient.</p>
<fig id="fig-14"><label>Figure 14</label><caption><title>Relation between area ratio of surface bugholes and CIB bughole scale</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-14.png"/></fig>
<p>According to the CIB classification and the regression analysis equation of the area ratio and CIB bughole scale, the recommended rating of bughole in this study is shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6"><label>Table 6</label><caption><title>Recommended rating of bughole </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Bughole rating</td>
<td align="left">1</td>
<td align="left">2</td>
<td align="left">3</td>
<td align="left">4</td>
<td align="left">5</td>
<td align="left">6</td>
<td align="left">7</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Area ratio</td>
<td align="left">(0&#x025;, 0.15&#x025;]</td>
<td align="left">(0.15&#x025;, 0.3&#x025;]</td>
<td align="left">(0.3&#x025;, 0.7&#x025;]</td>
<td align="left">(0.7&#x025;, 1.6&#x025;]</td>
<td align="left">(1.6&#x025;, 3&#x025;]</td>
<td align="left">(3&#x025;, 6&#x025;]</td>
<td align="left">(6&#x025;, 9&#x025;]</td>
</tr>
<tr>
<td align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The proposed method can output the pixels of each bughole, and the area ratio of surface bugholes is calculated by dividing the number of pixels in detected bugholes by the total number of pixels of the image. The final output layer is designed to directly output the calculated area ratio of the bugholes and a bughole rating by automatically comparing the recommended bughole scale shown in <xref ref-type="table" rid="table-6">Table 6</xref>. <xref ref-type="fig" rid="fig-15">Fig. 15</xref> illustrates some examples of bughole evaluation of concrete surfaces by the proposed method.</p>
<fig id="fig-15"><label>Figure 15</label><caption><title>Examples of bughole evaluation of concrete surfaces by the proposed method: (a) bughole rating is 2; (b) bughole rating is 3; (c) bughole rating is 4; (d) bughole rating is 5</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-15a.png"/><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-15b.png"/></fig>
</sec>
</sec>
<sec id="s_4">
<label>4</label>
<title>Discussion</title>
<p>The above analysis demonstrates that the evaluation method introduced in this study ensures better accuracy in detecting bugholes of concrete surfaces, which in turn provides better data for the improvement of concrete quality. The performance of the introduced bughole evaluation method is compared with that of the other three algorithms, namely Faster R-CNN, Retinanet, and FCN. The results are shown in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7"><label>Table 7</label><caption><title>Comparison of the detection performance between the proposed Mask R-CNN and the other three algorithms (Faster R-CNN, Retinanet, and FCN) </title></caption>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Algorithm</td>
<td align="center" colspan="2">Bounding-box AP (&#x025;)</td>
<td align="center" colspan="2">Segmentation AP (&#x025;)</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left"></td>
<td align="left">IoU = 0.5</td>
<td align="left">IoU = 0.75</td>
<td align="left">IoU = 0.5</td>
<td align="left">IoU = 0.75</td>
</tr>
<tr>
<td align="left"></td>
</tr>
<tr>
<td align="left">Faster R-CNN</td>
<td align="left">87.6</td>
<td align="left">65.4</td>
<td align="left">-</td>
<td align="left">-</td>
</tr>
<tr>
<td align="left">Retinanet</td>
<td align="left">82.5</td>
<td align="left">62.8</td>
<td align="left">-</td>
<td align="left">-</td>
</tr>
<tr>
<td align="left">Mask R-CNN</td>
<td align="left"><bold>90.0</bold></td>
<td align="left"><bold>67.5</bold></td>
<td align="left"><bold>90.8</bold></td>
<td align="left"><bold>64.7</bold></td>
</tr>
<tr>
<td align="left">FCN</td>
<td align="left">-</td>
<td align="left">-</td>
<td align="left">85.4</td>
<td align="left">61.2</td>
</tr>
<tr>
<td align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-7">Table 7</xref>, Faster R-CNN has a bounding box AP of 87.6&#x025; and 65.4&#x025; at IoU = 0.5 and IoU = 0.75, respectively. Retinanet has a bounding box AP of 82.5&#x025; and 62.8&#x025; at IoU = 0.5 and IoU = 0.75, respectively. Mask R-CNN has a bounding box AP of 90.0&#x025; and 67.5&#x025; at IoU = 0.5 and IoU = 0.75, respectively. To compare with the traditional CNN-based object detection algorithm, <xref ref-type="fig" rid="fig-16">Fig. 16</xref> shows the recognition results of the same bughole image under different algorithms. As shown in the results, bounding boxes can locate bugholes well, however, they are unable to identify the contours of the bugholes. Therefore, the surface bughole detection and evaluation methods based on object detection algorithms are not accurate enough. In contrast, Mask R-CNN outputs the instance segmentation mask while locating the bughole, which enables further quantification and evaluation. The test results in <xref ref-type="table" rid="table-7">Table 7</xref> show that the segmentation AP of Mask R-CNN at IoU = 0.5 and IoU = 0.75 is 90.8&#x025; and 64.7&#x025;, respectively. The segmentation AP of FCN at IoU = 0.5 and IoU = 0.75 is 85.4&#x025; and 61.2&#x025;, respectively. It is considered reasonable to choose the Mask R-CNN framework based on the instance segmentation algorithm for bughole identification and evaluation.</p>
<fig id="fig-16"><label>Figure 16</label><caption><title>Recognition results of Faster R-CNN and Mask R-CNN</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-16.png"/></fig>
<p>The application of the proposed method in detecting bugholes may still misidentify other surface defects as bugholes (i.e., false-positive results). For example, a spalling is misidentified as a bughole, as shown in the red box in <xref ref-type="fig" rid="fig-17">Fig. 17</xref>. Such false-positive results will affect the accuracy of concrete quality evaluation.</p>
<fig id="fig-17"><label>Figure 17</label><caption><title>Detection result of an example image containing bughole and spalling</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19082-fig-17.png"/></fig>
</sec>
<sec id="s_5">
<label>5</label>
<title>Conclusion</title>
<p>Concrete surface quality control is one of the keys in construction project management. The detection and evaluation of concrete surface bugholes are among the main tasks in implementing concrete quality control. Traditionally, project managers and inspectors mainly rely on human visual methods to evaluate the surface quality of concrete, which is subjective and inefficient. To mitigate these weaknesses, this paper has introduced a deep learning-based automatic instance-level evaluation method for detecting concrete surface bugholes more objectively and automatically.</p>
<p>The analysis in the paper suggests that Mask R-CNN-based method prevails over traditional methods for analyzing concrete quality in several aspects. Firstly, it improves the objectivity of the evaluation process. The method presented in this paper uses images as the basis for evaluation. A total of 273 concrete surface photos are collected to build a data set, and then image recognition technology based on deep learning replaces human eyes for detection and evaluation, which significantly improves the objectivity of evaluation. Secondly, the accuracy of detection is improved. The bughole detection model based on Mask R-CNN can identify the contours of bugholes effectively, and output the mask of each bughole instead of a bounding box so that the actual area of the bughole on the concrete surface can be calculated more accurately. The test results show that the average precision (AP) of bughole masks reaches 90.8&#x025;. Thirdly, the introduced method can provide better data for improving the quality of concrete. The method presented in this paper can recognize bugholes at an instance level and output the pixels of each bughole, and the area ratio of bugholes is automatically calculated, which in turn can provide better decision-making data about the quality of concrete surfaces.</p>
<p>Despite the excellent performance of the proposed method, it still has limitations in detecting bugholes due to the interference of other imperfections, such as spalling. It is recommended to develop detection and quantification methods for multiple damage types based on deep learning to establish a more comprehensive evaluation model for concrete exterior quality.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> This work is supported by Chongqing Municipal Natural Science Foundation (Grant Nos. cstc2021jcyj-bsh0189 and cstc2019jcyj-bshX0070) and Chongqing Jiulongpo District Science and Technology Planning Project (Grant No. 2020&#x2013;01&#x2013;001-Y).</p></fn>
<fn fn-type="other"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Ding</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Luo</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>G.</given-names></string-name> et al.</person-group> (<year>2021</year>). <article-title>Automated bughole detection and quality performance assessment of concrete using image processing and deep convolutional neural networks</article-title>. <source>Construction and Building Materials</source><italic>,</italic> <volume>281</volume><italic>,</italic> <fpage>122576</fpage>. DOI <pub-id pub-id-type="doi">10.1016/j.conbuildmat.2021.122576</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yoshitake</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Maeda</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Hieda</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Image analysis for the detection and quantification of concrete bugholes in a tunnel lining</article-title>. <source>Case Studies in Construction Materials</source><italic>,</italic> <volume>8</volume><italic>,</italic> <fpage>116</fpage>&#x2013;<lpage>130</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.cscm.2018.01.002</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kalayci</surname>, <given-names>A. S.</given-names></string-name>, <string-name><surname>Yalim</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Mirmiran</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2009</year>). <article-title>Effect of untreated surface disbonds on performance of FRP-retrofitted concrete beams</article-title>. <source>Journal of Composites for Construction</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>476</fpage>&#x2013;<lpage>485</lpage>. DOI <pub-id pub-id-type="doi">10.1061/(asce)cc.1943-5614.0000032</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ichimiya</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Idemitsu</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Yamasaki</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2002</year>). <article-title>The influence of air content and fluidity of mortar on the characteristics of surface voids in self-compacting concrete</article-title>. <source>Doboku Gakkai Ronbunshu</source><italic>,</italic> <volume>56</volume><italic>(</italic><issue>711</issue><italic>),</italic> <fpage>135</fpage>&#x2013;<lpage>146</lpage>. DOI <pub-id pub-id-type="doi">10.2208/jscej.2002.711_135</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Yao</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Instance-level recognition and quantification for concrete surface bughole based on deep learning</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>107</volume><italic>(</italic><issue>174</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>13</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.autcon.2019.102920</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lemaire</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Escadeillas</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Ringot</surname>, <given-names>E.</given-names></string-name></person-group> (<year>2005</year>). <article-title>Evaluating concrete surfaces using an image analysis process</article-title>. <source>Construction and Building Materials</source><italic>,</italic> <volume>19</volume><italic>,</italic> <fpage>604</fpage>&#x2013;<lpage>611</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.conbuildmat.2005.01.025</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Coster</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Chermant</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2001</year>). <article-title>Image analysis and mathematical morphology for civil engineering materials</article-title>. <source>Cement and Concrete Composites</source><italic>,</italic> <volume>23</volume><italic>,</italic> <fpage>133</fpage>&#x2013;<lpage>151</lpage>. DOI <pub-id pub-id-type="doi">10.1016/S0958-9465(00)00058-5</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chermant</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2001</year>). <article-title>Why automatic image analysis ? An introduction to this issue</article-title>. <source>Cement &#x0026; Concrete Composites</source><italic>,</italic> <volume>23</volume><italic>,</italic> <fpage>127</fpage>&#x2013;<lpage>131</lpage>. DOI <pub-id pub-id-type="doi">10.1016/S0958-9465(00)00077-9</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Chang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Skibniewski</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2006</year>). <article-title>Automated recognition of surface defects using digital color image processing</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>15</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>540</fpage>&#x2013;<lpage>549</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.autcon.2005.08.001</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Image analysis for detection of bugholes on concrete surface</article-title>. <source>Construction and Building Materials</source><italic>,</italic> <volume>137</volume><italic>,</italic> <fpage>432</fpage>&#x2013;<lpage>440</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.conbuildmat.2017.01.098</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Laofor</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Peansupap</surname>, <given-names>V.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Defect detection and quantification system to support subjective visual quality inspection via a digital image processing: A tiling work case study</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>24</volume><italic>,</italic> <fpage>160</fpage>&#x2013;<lpage>174</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.autcon.2012.02.012</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Silva</surname>, <given-names>W. R. L. D.</given-names></string-name>, <string-name><surname>&#352;temberk</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Expert system applied for classifying self-compacting concrete surface finish</article-title>. <source>Advances in Engineering Software</source><italic>,</italic> <volume>64</volume><italic>,</italic> <fpage>47</fpage>&#x2013;<lpage>61</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.advengsoft.2013.04.005</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ince</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Unsupervised classification of polarimetric SAR image with dynamic clustering: An image processing approach</article-title>. <source>Advances in Engineering Software</source><italic>,</italic> <volume>41</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>636</fpage>&#x2013;<lpage>646</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.advengsoft.2009.12.004</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ozkul</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Kucuk</surname>, <given-names>I.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Design and optimization of an instrument for measuring bughole rating of concrete surfaces</article-title>. <source>Journal of the Franklin Institute</source><italic>,</italic> <volume>348</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>1377</fpage>&#x2013;<lpage>1392</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.jfranklin.2010.04.004</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khireddine</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Benmahammed</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Puech</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2007</year>). <article-title>Digital image restoration by wiener filter in 2D case</article-title>. <source>Advances in Engineering Software</source><italic>,</italic> <volume>38</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>513</fpage>&#x2013;<lpage>516</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.advengsoft.2006.10.001</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yao</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Deep-learning-based bughole detection for concrete surface image</article-title>. <source>Advances in Civil Engineering</source><italic>,</italic> <volume>2019</volume>. DOI <pub-id pub-id-type="doi">10.1155/2019/8582963</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sutskever</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Hinton</surname>, <given-names>G. E.</given-names></string-name></person-group> (<year>2012</year>). <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Communications of the ACM</source><italic>,</italic> <volume>60</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>84</fpage>&#x2013;<lpage>90</lpage>. DOI <pub-id pub-id-type="doi">10.1145/3065386</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Deep learning: Methods and applications</article-title>. <source>Foundations and Trends in Signal Processing</source><italic>,</italic> <volume>7</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>197</fpage>&#x2013;<lpage>387</lpage>. DOI <pub-id pub-id-type="doi">10.1561/2000000039</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tedeschi</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Benedetto</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A Real-time automatic pavement crack and pothole recognition system for mobile android-based devices</article-title>. <source>Advanced Engineering Informatics</source><italic>,</italic> <volume>32</volume><italic>,</italic> <fpage>11</fpage>&#x2013;<lpage>25</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.aei.2016.12.004</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cha</surname>, <given-names>Y. J.</given-names></string-name>, <string-name><surname>Choi</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Suh</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Mahmoudkhani</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>B&#x00FC;y&#x00FC;k&#x00F6;zt&#x00FC;rk</surname>, <given-names>O.</given-names></string-name></person-group> (<year>2017b</year>). <article-title>Autonomous structural visual inspection using region-based deep learning for detecting multiple damage types</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>33</volume><italic>(</italic><issue>9</issue><italic>),</italic> <fpage>731</fpage>&#x2013;<lpage>747</lpage>. DOI <pub-id pub-id-type="doi">10.1111/mice.12334</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>F. C.</given-names></string-name>, <string-name><surname>Jahanshahi</surname>, <given-names>M. R.</given-names></string-name></person-group> (<year>2018</year>). <article-title>NB-CNN: Deep learning-based crack detection using convolutional neural network and naive Bayes data fusion</article-title>. <source>IEEE Transactions on Industrial Electronics</source><italic>,</italic> <volume>65</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>4392</fpage>&#x2013;<lpage>4400</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TIE.2017.2764844</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>K. C. P.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>J. Q.</given-names></string-name>, <string-name><surname>Fei</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Deep learning for asphalt pavement cracking recognition using convolutional neural network</article-title>. <conf-name>International Conference on Highway Pavements and Airfield Technology</conf-name>, pp. <fpage>166</fpage>&#x2013;<lpage>177</lpage>. <publisher-loc>Philadelphia, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cire&#351;an</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Meier</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Masci</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Schmidhuber</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Multi-column deep neural network for traffic sign classification</article-title>. <source>Neural Networks</source><italic>,</italic> <volume>32</volume><italic>,</italic> <fpage>333</fpage>&#x2013;<lpage>338</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.neunet.2012.02.023</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cha</surname>, <given-names>Y. J.</given-names></string-name>, <string-name><surname>Choi</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>B&#x00FC;y&#x00FC;k&#x00F6;zt&#x00FC;rk</surname>, <given-names>O.</given-names></string-name></person-group> (<year>2017a</year>). <article-title>Deep learning-based crack damage detection using convolutional neural networks</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>32</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>361</fpage>&#x2013;<lpage>378</lpage>. DOI <pub-id pub-id-type="doi">10.1111/mice.12263</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cui</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Dai</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Intelligent recognition of erosion damage to concrete based on improved YOLO-v3</article-title>. <source>Materials Letters</source><italic>,</italic> <volume>302</volume><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>4</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.matlet.2021.130363</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shin</surname>, <given-names>H. K.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>S. W.</given-names></string-name>, <string-name><surname>Hong</surname>, <given-names>G. P.</given-names></string-name>, <string-name><surname>Sael</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>S. H.</given-names></string-name> et al.</person-group> (<year>2020</year>). <article-title>Defect-detection model for underground parking lots using image object-detection method</article-title>. <source>Computers, Materials and Continua</source><italic>,</italic> <volume>66</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>2493</fpage>&#x2013;<lpage>2507</lpage>. DOI <pub-id pub-id-type="doi">10.32604/cmc.2021.014170</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Atha</surname>, <given-names>D. J.</given-names></string-name>, <string-name><surname>Jahanshahi</surname>, <given-names>M. R.</given-names></string-name></person-group>, (<year>2018</year>). <article-title>Evaluation of deep learning approaches based on convolutional neural networks for corrosion detection</article-title>. <source>Structural Health Monitoring</source><italic>,</italic> <volume>17</volume><italic>,</italic> <fpage>1110</fpage>&#x2013;<lpage>1128</lpage>. DOI <pub-id pub-id-type="doi">10.1177/1475921717737051</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koch</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Georgieva</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Kasireddy</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Akinci</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Fieguth</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2015</year>). <article-title>A review on computer vision based defect detection and condition assessment of concrete and asphalt civil infrastructure</article-title>. <source>Advanced Engineering Informatics</source><italic>,</italic> <volume>29</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>196</fpage>&#x2013;<lpage>210</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.aei.2015.01.008</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Gkioxari</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Dollar</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Mask R-CNN</article-title>. <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>, pp. <fpage>2980</fpage><publisher-loc>&#x2013;</publisher-loc><lpage>2988</lpage><publisher-loc>. Venice, Italy</publisher-loc>.</mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Pixel-level automatic annotation for forest fire image</article-title>. <source>Engineering Applications of Artificial Intelligence</source><italic>,</italic> <volume>104</volume><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>14</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.engappai.2021.104353</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhagat</surname>, <given-names>P. K.</given-names></string-name>, <string-name><surname>Choudhary</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Image annotation: then and now</article-title>. <source>Image and Vision Computing</source><italic>,</italic> <volume>80</volume><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>23</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.imavis.2018.09.017</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. <conf-name>IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>770</fpage><publisher-loc>&#x2013;</publisher-loc><lpage>778</lpage><publisher-loc>. Las Vegas, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>T. Y.</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Hariharan</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Feature pyramid networks for object detection</article-title>. <conf-name>30th IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>936</fpage>&#x2013;<lpage>944</lpage>. <publisher-loc>Hawaii, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Faster R-CNN: Towards real-time object detection with region proposal networks</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source><italic>,</italic> <volume>39</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1137</fpage>&#x2013;<lpage>1149</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shelhamer</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Long</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Darrell</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Fully convolutional networks for semantic segmentation</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source><italic>,</italic> <volume>39</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>640</fpage>&#x2013;<lpage>651</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2572683</pub-id>.</mixed-citation></ref>
</ref-list>
<sec id="s_6_1">
<label>6.1</label>
<title>Appendix A: The APs of both bounding box and mask</title>
<table-wrap id="table-8"><label>Table 8</label>
<table>
<colgroup>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
<col valign="top" align="left"/>
</colgroup>
<thead valign="bottom">
<tr>
<td align="left">Image</td>
<td align="left" colspan="2">Bounding-box (&#x025;)</td>
<td align="left" colspan="2">Mask (&#x025;)</td>
</tr>
<tr>
<td align="left"></td>
<td align="left">IoU = 0.5</td>
<td align="left">IoU = 0.75</td>
<td align="left">mAP</td>
<td align="left">IoU = 0.5</td>
<td align="left">IoU = 0.75</td>
<td align="left">mAP</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">0.972</td>
<td align="left">0.834</td>
<td align="left">0.903</td>
<td align="left">0.972</td>
<td align="left">0.580</td>
<td align="left">0.776</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">0.927</td>
<td align="left">0.866</td>
<td align="left">0.897</td>
<td align="left">0.927</td>
<td align="left">0.526</td>
<td align="left">0.727</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">0.956</td>
<td align="left">0.849</td>
<td align="left">0.903</td>
<td align="left">0.956</td>
<td align="left">0.696</td>
<td align="left">0.826</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">0.983</td>
<td align="left">0.888</td>
<td align="left">0.936</td>
<td align="left">0.956</td>
<td align="left">0.804</td>
<td align="left">0.880</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">0.858</td>
<td align="left">0.757</td>
<td align="left">0.808</td>
<td align="left">0.858</td>
<td align="left">0.641</td>
<td align="left">0.750</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">0.938</td>
<td align="left">0.800</td>
<td align="left">0.869</td>
<td align="left">0.938</td>
<td align="left">0.664</td>
<td align="left">0.801</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">0.974</td>
<td align="left">0.934</td>
<td align="left">0.954</td>
<td align="left">0.974</td>
<td align="left">0.896</td>
<td align="left">0.935</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">0.958</td>
<td align="left">0.833</td>
<td align="left">0.896</td>
<td align="left">0.958</td>
<td align="left">0.704</td>
<td align="left">0.831</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">0.853</td>
<td align="left">0.764</td>
<td align="left">0.809</td>
<td align="left">0.946</td>
<td align="left">0.667</td>
<td align="left">0.807</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">0.961</td>
<td align="left">0.756</td>
<td align="left">0.859</td>
<td align="left">0.983</td>
<td align="left">0.434</td>
<td align="left">0.709</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">0.946</td>
<td align="left">0.836</td>
<td align="left">0.891</td>
<td align="left">0.906</td>
<td align="left">0.420</td>
<td align="left">0.663</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">0.863</td>
<td align="left">0.750</td>
<td align="left">0.807</td>
<td align="left">0.863</td>
<td align="left">0.407</td>
<td align="left">0.635</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">0.963</td>
<td align="left">0.826</td>
<td align="left">0.895</td>
<td align="left">0.963</td>
<td align="left">0.472</td>
<td align="left">0.718</td>
</tr>
<tr>
<td align="left">14</td>
<td align="left">0.889</td>
<td align="left">0.839</td>
<td align="left">0.864</td>
<td align="left">0.889</td>
<td align="left">0.685</td>
<td align="left">0.787</td>
</tr>
<tr>
<td align="left">15</td>
<td align="left">0.972</td>
<td align="left">0.799</td>
<td align="left">0.886</td>
<td align="left">0.972</td>
<td align="left">0.799</td>
<td align="left">0.886</td>
</tr>
<tr>
<td align="left">16</td>
<td align="left">0.942</td>
<td align="left">0.902</td>
<td align="left">0.922</td>
<td align="left">0.942</td>
<td align="left">0.902</td>
<td align="left">0.922</td>
</tr>
<tr>
<td align="left">17</td>
<td align="left">0.985</td>
<td align="left">0.965</td>
<td align="left">0.975</td>
<td align="left">0.985</td>
<td align="left">0.901</td>
<td align="left">0.943</td>
</tr>
<tr>
<td align="left">18</td>
<td align="left">0.914</td>
<td align="left">0.829</td>
<td align="left">0.872</td>
<td align="left">0.870</td>
<td align="left">0.440</td>
<td align="left">0.665</td>
</tr>
<tr>
<td align="left">19</td>
<td align="left">0.834</td>
<td align="left">0.712</td>
<td align="left">0.773</td>
<td align="left">0.804</td>
<td align="left">0.746</td>
<td align="left">0.775</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left">0.968</td>
<td align="left">0.908</td>
<td align="left">0.938</td>
<td align="left">0.968</td>
<td align="left">0.762</td>
<td align="left">0.865</td>
</tr>
<tr>
<td align="left">21</td>
<td align="left">0.805</td>
<td align="left">0.732</td>
<td align="left">0.769</td>
<td align="left">0.731</td>
<td align="left">0.586</td>
<td align="left">0.659</td>
</tr>
<tr>
<td align="left">22</td>
<td align="left">0.896</td>
<td align="left">0.866</td>
<td align="left">0.881</td>
<td align="left">0.896</td>
<td align="left">0.794</td>
<td align="left">0.845</td>
</tr>
<tr>
<td align="left">23</td>
<td align="left">0.876</td>
<td align="left">0.697</td>
<td align="left">0.787</td>
<td align="left">0.876</td>
<td align="left">0.346</td>
<td align="left">0.611</td>
</tr>
<tr>
<td align="left">24</td>
<td align="left">0.957</td>
<td align="left">0.917</td>
<td align="left">0.937</td>
<td align="left">0.957</td>
<td align="left">0.897</td>
<td align="left">0.927</td>
</tr>
<tr>
<td align="left">25</td>
<td align="left">0.963</td>
<td align="left">0.862</td>
<td align="left">0.913</td>
<td align="left">0.963</td>
<td align="left">0.863</td>
<td align="left">0.913</td>
</tr>
<tr>
<td align="left">26</td>
<td align="left">0.968</td>
<td align="left">0.909</td>
<td align="left">0.939</td>
<td align="left">0.988</td>
<td align="left">0.632</td>
<td align="left">0.810</td>
</tr>
<tr>
<td align="left">27</td>
<td align="left">0.958</td>
<td align="left">0.898</td>
<td align="left">0.928</td>
<td align="left">0.958</td>
<td align="left">0.771</td>
<td align="left">0.865</td>
</tr>
<tr>
<td align="left">28</td>
<td align="left">0.897</td>
<td align="left">0.743</td>
<td align="left">0.820</td>
<td align="left">0.897</td>
<td align="left">0.695</td>
<td align="left">0.796</td>
</tr>
<tr>
<td align="left">29</td>
<td align="left">0.859</td>
<td align="left">0.759</td>
<td align="left">0.809</td>
<td align="left">0.759</td>
<td align="left">0.759</td>
<td align="left">0.759</td>
</tr>
<tr>
<td align="left">30</td>
<td align="left">0.919</td>
<td align="left">0.735</td>
<td align="left">0.827</td>
<td align="left">0.919</td>
<td align="left">0.383</td>
<td align="left">0.651</td>
</tr>
<tr>
<td align="left">31</td>
<td align="left">0.976</td>
<td align="left">0.926</td>
<td align="left">0.951</td>
<td align="left">0.976</td>
<td align="left">0.903</td>
<td align="left">0.940</td>
</tr>
<tr>
<td align="left">32</td>
<td align="left">0.916</td>
<td align="left">0.730</td>
<td align="left">0.823</td>
<td align="left">0.916</td>
<td align="left">0.397</td>
<td align="left">0.657</td>
</tr>
<tr>
<td align="left">33</td>
<td align="left">0.855</td>
<td align="left">0.815</td>
<td align="left">0.835</td>
<td align="left">0.855</td>
<td align="left">0.805</td>
<td align="left">0.830</td>
</tr>
<tr>
<td align="left">34</td>
<td align="left">0.922</td>
<td align="left">0.720</td>
<td align="left">0.821</td>
<td align="left">0.922</td>
<td align="left">0.720</td>
<td align="left">0.821</td>
</tr>
<tr>
<td align="left">35</td>
<td align="left">0.933</td>
<td align="left">0.819</td>
<td align="left">0.876</td>
<td align="left">0.933</td>
<td align="left">0.769</td>
<td align="left">0.851</td>
</tr>
<tr>
<td align="left">36</td>
<td align="left">0.966</td>
<td align="left">0.770</td>
<td align="left">0.868</td>
<td align="left">0.966</td>
<td align="left">0.765</td>
<td align="left">0.866</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</back>
</article>