<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">SDHM</journal-id>
<journal-id journal-id-type="nlm-ta">SDHM</journal-id>
<journal-id journal-id-type="publisher-id">SDHM</journal-id>
<journal-title-group>
<journal-title>Structural Durability &#x0026; Health Monitoring</journal-title>
</journal-title-group>
<issn pub-type="epub">1930-2991</issn>
<issn pub-type="ppub">1930-2983</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">42388</article-id>
<article-id pub-id-type="doi">10.32604/sdhm.2024.042388</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Rapid and Accurate Identification of Concrete Surface Cracks via a Lightweight &#x0026; Efficient YOLOv3 Algorithm</article-title><alt-title alt-title-type="left-running-head">Rapid and Accurate Identification of Concrete Surface Cracks via a Lightweight &#x0026; Efficient YOLOv3 Algorithm</alt-title><alt-title alt-title-type="right-running-head">Rapid and Accurate Identification of Concrete Surface Cracks via a Lightweight &#x0026; Efficient YOLOv3 Algorithm</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Gu</surname><given-names>Haoan</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Zhu</surname><given-names>Kai</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Strauss</surname><given-names>Alfred</given-names></name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Shi</surname><given-names>Yehui</given-names></name>
<xref ref-type="aff" rid="aff-3">3</xref>
<xref ref-type="aff" rid="aff-4">4</xref>
</contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Sumarac</surname><given-names>Dragoslav</given-names></name>
<xref ref-type="aff" rid="aff-5">5</xref>
</contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Cao</surname><given-names>Maosen</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref><email>cmszhy@hhu.edu.cn</email>
</contrib>
<aff id="aff-1"><label>1</label><institution>College of Mechanics and Engineering Science, Hohai University</institution>, <addr-line>Nanjing, 211100</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Civil Engineering and Natural Hazards, University of Natural Resources and Life Sciences</institution>, <addr-line>Vienna, 1180</addr-line>, <country>Austria</country></aff>
<aff id="aff-3"><label>3</label><institution>The First Geological Brigade of the Bureau of Geology and Mineral Resources of Jiangsu</institution>, <addr-line>Nanjing, 210041</addr-line>, <country>China</country></aff>
<aff id="aff-4"><label>4</label><institution>Control Technology Group Co.</institution>, <addr-line>Nanjing, 210041</addr-line>, <country>China</country></aff>
<aff id="aff-5"><label>5</label><institution>Department of Technical Sciences, Civil Engineering, State University of Novi Pazar</institution>, <addr-line>Novi Pazar, 36300</addr-line>, <country>Serbia</country></aff>
</contrib-group><author-notes><corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Maosen Cao. Email: <email>cmszhy@hhu.edu.cn</email></corresp></author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>05</day><month>6</month><year>2024</year></pub-date>
<volume>18</volume>
<issue>4</issue>
<fpage>363</fpage>
<lpage>380</lpage>
<history>
<date date-type="received"><day>28</day><month>5</month><year>2023</year></date>
<date date-type="accepted"><day>01</day><month>11</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Gu et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Gu et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="_SDHM_42388.pdf"></self-uri>
<abstract>
<p>Concrete materials and structures are extensively used in transformation infrastructure and they usually bear cracks during their long-term operation. Detecting cracks using deep-learning algorithms like YOLOv3 (You Only Look Once version 3) is a new trend to pursue intelligent detection of concrete surface cracks. YOLOv3 is a typical deep-learning algorithm used for object detection. Owing to its generality, YOLOv3 lacks specific efficiency and accuracy in identifying concrete surface cracks. An improved algorithm based on YOLOv3, specialized in the rapid and accurate identification of concrete surface cracks is worthy of investigation. This study proposes a tailored deep-learning algorithm, termed MDN-YOLOv3 (MDN: multi-dilated network), of which the MDN is formulated based on three retrofit techniques, and it provides a new backbone network for YOLOv3. The three specific retrofit techniques are briefed: (i) Depthwise separable convolution is utilized to reduce the size of the backbone network; (ii) The dilated-down sampling structure is proposed and used in the backbone network to achieve multi-scale feature fusion; and (iii) The convolutional block attention module is introduced to enhance feature extraction ability. Results show that the proposed MDN-YOLOv3 is 97.2% smaller and 41.5% faster than YOLOv3 in identifying concrete surface cracks, forming a lightweight &#x0026; efficient YOLOv3 algorithm for intelligently identifying concrete surface cracks.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>YOLOv3</kwd>
<kwd>concrete crack identification</kwd>
<kwd>MDN-YOLOv3</kwd>
<kwd>optimization retrofit techniques</kwd>
<kwd>depthwise separable convolution</kwd>
<kwd>dilated-down sampling</kwd>
<kwd>attention mechanism</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>International Science &#x0026; Technology Cooperation Project of Jiangsu Province</funding-source>
<award-id>BZ2022010</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Jiangsu-Czech Bilateral Co-Funding R&#x0026;D Project</funding-source>
<award-id>BZ2023011</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Innovative Talents Exchange Foreign Experts Project</funding-source>
<award-id>DL2023019001L</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Concrete materials and structures have been extensively used in transportation infrastructure, typically highways and railways [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. These materials and structures inevitably bear cracks during their long-term operation [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Owing to the large-scale characteristic of transportation infrastructure, the distribution of cracks in concrete materials and structures often exhibits divergence and dispersion [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. The occurrence of cracks may impair the integrity and performance of related infrastructural systems [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. The rapid and accurate detection of cracks in concrete materials and structures becomes exceptionally crucial to the safety of transformation infrastructure [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>]. Therefore, it is of great significance to develop an efficient technique to rapidly identify crack locations and patterns and evaluate the service performance of concrete structures [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>Crack identification currently mainly relies on conventional routine-based visual inspection [<xref ref-type="bibr" rid="ref-16">16</xref>]. Unfortunately, visual inspection-based condition assessment is time-consuming and depends on the experience and knowledge of inspectors [<xref ref-type="bibr" rid="ref-17">17</xref>]. In general, visual inspection cannot authentically discover abrupt and subtle cracks and may miss potential risks in concrete structures [<xref ref-type="bibr" rid="ref-18">18</xref>]. To address this issue, computer vision-aided intelligent identification methods have become a research focus for the online detection of surface cracks. Numerous image processing methods have been utilized to detect surface defects or damage in structures. The main implementation methods can be divided into three categories, namely, digital image processing (DIP), machine learning (ML), and deep learning (DL) methods.</p>
<p>DIP methods usually conduct edge detection or pattern recognition of images to identify cracks or defects based on the abnormality of pixel points [<xref ref-type="bibr" rid="ref-19">19</xref>]. The locations of cracks or defects can be identified with advanced 2D signal processing techniques such as various types of transformations [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>] and filtering methods [<xref ref-type="bibr" rid="ref-22">22</xref>]. In addition, DIP methods can identify the crack width when combined with the Laplacian method [<xref ref-type="bibr" rid="ref-23">23</xref>]. Many environmental or artificial interfering factors such as lighting, background layers, noise, and threshold setting may significantly impact the identified results [<xref ref-type="bibr" rid="ref-24">24</xref>&#x2013;<xref ref-type="bibr" rid="ref-27">27</xref>]. However, ML algorithms can detect cracks or defects with numerous applications in practice [<xref ref-type="bibr" rid="ref-28">28</xref>&#x2013;<xref ref-type="bibr" rid="ref-32">32</xref>]. The algorithm can accomplish crack identification and classification integrated with DIP methods, but the generality of algorithms needs further expansion for low-resolution images with background and complex multi-classification problems.</p>
<p>DL methods can effectively extract image characteristics with environmental interference and are especially powerful in processing large-scale training datasets and multi-classification problems compared with DIP and ML methods. Convolutional neural networks (CNNs), a notable method in DL, have garnered widespread utilization in the realms of crack identification due to their superior multi-scale feature extraction, noise resilience, and broad-spectrum recognition capacities [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-35">35</xref>]. Crack identification currently encompasses both object detection [<xref ref-type="bibr" rid="ref-36">36</xref>&#x2013;<xref ref-type="bibr" rid="ref-38">38</xref>] and semantic segmentation [<xref ref-type="bibr" rid="ref-39">39</xref>&#x2013;<xref ref-type="bibr" rid="ref-42">42</xref>]. For example, an initial CNN can detect whether there are cracks in pictures, which is a dichotomous classification problem [<xref ref-type="bibr" rid="ref-43">43</xref>]. However, an initial CNN cannot identify crack locations and is disabled to take a further assessment of structures. To address this issue, a multi-layered image preprocessing strategy has been proposed to locate cracks in concrete structures in the measured pictures [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>]. The cracking region is filled up with a large number of bounding boxes, which are not flexible enough to directly determine the region of cracks [<xref ref-type="bibr" rid="ref-46">46</xref>]. Further, several two-stage identification methods, including regional-based CNNs [<xref ref-type="bibr" rid="ref-47">47</xref>] and faster regional-based CNNs [<xref ref-type="bibr" rid="ref-48">48</xref>], were developed for object detection in images. The crack region can be first identified and then classified based on the above two-stage methods. Additionally, crack width can be identified based on improved mask-RCNN methods [<xref ref-type="bibr" rid="ref-49">49</xref>]. You Only Look Once (YOLO) [<xref ref-type="bibr" rid="ref-50">50</xref>] was recently proposed to accomplish the one-stage identification of location and classification of surface cracks or defects with applications in bridges [<xref ref-type="bibr" rid="ref-51">51</xref>] and pavement [<xref ref-type="bibr" rid="ref-52">52</xref>&#x2013;<xref ref-type="bibr" rid="ref-54">54</xref>] or tunnel engineering [<xref ref-type="bibr" rid="ref-55">55</xref>] and aircraft structures [<xref ref-type="bibr" rid="ref-56">56</xref>]. The architectures of models such as YOLOv3, other single-stage and dual-stage detection networks, and semantic segmentation frameworks [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>] are notably intricate. These models, characterized by their voluminous parameter sets, mandate exhaustive training processes. As a result, ensuring precise identification outcomes demands not only potent computational prowess but also ample memory capacity [<xref ref-type="bibr" rid="ref-51">51</xref>]. A simple network structure is necessary to improve computation speed and reduce memory consumption in order to implement the CNN algorithm on mobile terminals or unmanned aerial vehicles. A series of YOLO-based networks with fewer training coefficients, such as MobileNetv2-YOLOv3 (M2-YOLOv3) [<xref ref-type="bibr" rid="ref-57">57</xref>] and MobileNetv3-YOLOv3 (M3-YOLOv3) [<xref ref-type="bibr" rid="ref-58">58</xref>], were proposed to achieve the purpose of real-time object detection in mobile terminals [<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>]. These networks are more efficient and designed to detect hundreds of different objects; however, there are only tens of concrete crack classes. Therefore, the YOLO algorithm still has room for optimization in concrete crack identification.</p>
<p>To address this deficiency, this study proposes a lightweight YOLOv3, called MDN-YOLOv3, for intelligently identifying the cracks of concrete structures using a self-built training set. The rest of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> briefly summarizes the structures and imperfections of YOLOv3 in identifying cracks in concrete structures. <xref ref-type="sec" rid="s3">Section 3</xref> proposes the improved design methods of the backbone network in MDN-YOLOv3. <xref ref-type="sec" rid="s4">Section 4</xref> compares the performance of the concrete crack identification of MDN-YOLOv3 with existing algorithms. <xref ref-type="sec" rid="s5">Section 5</xref> compares the test results and the accuracy of the method with existing models. <xref ref-type="sec" rid="s6">Section 6</xref> concludes and provides remarks on this study.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Inadequacies of YOLOv3 in Identifying Concrete Cracks</title>
<sec id="s2_1">
<label>2.1</label>
<title>Overview</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates the structure of YOLOv3, composed primarily of the Darknet-53 backbone and the feature pyramid network (FPN). The main identification process of YOLOv3 can be described as follows:</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The structure of YOLOv3</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-1.tif"/>
</fig>
<p>(1) Input process: The size of input images is adjusted to 416 &#x00D7; 416 &#x00D7; 3 for the feature extraction process.</p>
<p>(2) Feature extraction process: First, the images with 416 &#x00D7; 416 &#x00D7; 3 pixels are input into the backbone network of Darknet-53 [<xref ref-type="bibr" rid="ref-61">61</xref>] for feature extraction, and three effective feature layers with sizes of 13 &#x00D7; 13, 26 &#x00D7; 26, and 52 &#x00D7; 52 are output, as visualized in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The depth of the feature maps starts from the three channels of the RGB image and expands through subsequent layers with 32, 64, 128, 256, 512, and finally 1024 filters. Every layer is responsible for learning more complex features. Then, the three effective feature layers are input into the FPN. Through a series of operations, including continuous convolution operation, up-sampling, down-sampling, and feature fusion from the three effective feature layers, the enhanced feature maps of FPN are output.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Network details of YOLOv3. Convolutional <inline-formula id="ieqn-1">
<mml:math id="mml-ieqn-1"><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math>
</inline-formula> &#x003D; conv2d <inline-formula id="ieqn-2">
<mml:math id="mml-ieqn-2"><mml:mi>k</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:math>
</inline-formula> &#x002B; Nonlinear</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-2.tif"/>
</fig>
<p>(3) Output process: The output feature maps of FPN are used to predict three bounding boxes in every grid cell. The coordinate of the bounding box is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. Every bounding box predicts three key parameters based on extracted and multi-scale features: (i) Coordinate the dataset of the rectangular box, including horizontal coordinates <inline-formula id="ieqn-3">
<mml:math id="mml-ieqn-3"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>x</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> and vertical coordinates <inline-formula id="ieqn-4">
<mml:math id="mml-ieqn-4"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> of the centroid, width <inline-formula id="ieqn-5">
<mml:math id="mml-ieqn-5"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>w</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula>, and height <inline-formula id="ieqn-6">
<mml:math id="mml-ieqn-6"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:math>
</inline-formula> of the bounding box; (ii) Classification of detected objects; (iii) Confidence of predicted results.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Anchor box to predict box process</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-3.tif"/>
</fig>
<p>The loss function can be calculated as follows:<disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left" rowspacing=".5em" columnspacing="thickmathspace" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msqrt><mml:msubsup><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msubsup><mml:mrow><mml:mover><mml:mi>w</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:msqrt></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msqrt><mml:msubsup><mml:mi>h</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:msqrt><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:msqrt></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">b</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover><mml:mi>P</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>P</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-7">
<mml:math id="mml-ieqn-7"><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> and <inline-formula id="ieqn-8">
<mml:math id="mml-ieqn-8"><mml:mrow><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">b</mml:mi><mml:mi mathvariant="normal">j</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> are weight coefficients and remain unchanged in the loss function. <inline-formula id="ieqn-9">
<mml:math id="mml-ieqn-9"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> is the responsible anchor box that has the highest intersection over union (IoU) with the ground-truth box, which is actually used for predictions in the loss function. <inline-formula id="ieqn-10">
<mml:math id="mml-ieqn-10"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> represents the complementary set of responsible anchor boxes <inline-formula id="ieqn-11">
<mml:math id="mml-ieqn-11"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula>. <inline-formula id="ieqn-12">
<mml:math id="mml-ieqn-12"><mml:mi>S</mml:mi></mml:math>
</inline-formula> is the size of grid cells, and <inline-formula id="ieqn-13">
<mml:math id="mml-ieqn-13"><mml:mi>B</mml:mi></mml:math>
</inline-formula> is the number of anchor boxes. <inline-formula id="ieqn-14">
<mml:math id="mml-ieqn-14"><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:math>
</inline-formula> and <inline-formula id="ieqn-15">
<mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:math>
</inline-formula> are the confidence and probability of belonging to a specific category, respectively. The superscript represents the prediction center of the corresponding parameters.</p>
<p>The IoU can be calculated as follows:<disp-formula id="eqn-2"><label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2229;</mml:mo><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>&#x222A;</mml:mo><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:mstyle></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-16">
<mml:math id="mml-ieqn-16"><mml:mi>I</mml:mi><mml:mi>O</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</inline-formula> is the intersection over the union of the ground-truth boxes and anchor boxes.</p>
<p>The loss function serves to adjust the generated anchor box. It is primarily composed of three components: the loss associated with the box&#x2019;s coordinates, the confidence, and the category. By minimizing this loss function, YOLOv3 can accurately pinpoint the location of objects in every output image.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Inadequacies</title>
<p>YOLOv3 can detect multiple targets in a single forward pass, simultaneously predicting class scores and bounding boxes, which streamlines the object detection process. As a result, YOLOv3 is efficient and has found applications in various domains [<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>]. There are a series of CNNs used as the backbone networks of YOLO with fewer training coefficients, such as MobileNetv2 and MobileNetv3; however, the backbone networks in YOLO still do not meet the demand for swift crack identification with high accuracy. This is because the CNNs serving as the backbone network of YOLO in concrete crack identification are made for identifying hundreds of classes, whereas there are only tens of classes in concrete cracks. Therefore, the series of universal YOLOv3 models are large in size and time-consuming in identifying concrete cracks. In other words, the series of universal YOLOv3 models is relatively complex and unsuitable for concrete crack identification [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>,<xref ref-type="bibr" rid="ref-64">64</xref>]. To this end, a tailored deep-learning algorithm termed MDN-YOLOv3 is proposed.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>MDN-YOLOv3</title>
<sec id="s3_1">
<label>3.1</label>
<title>Retrofit Techniques</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Depthwise Separable Convolution (DSC)</title>
<p>DSC was proposed by Sandler et al. [<xref ref-type="bibr" rid="ref-60">60</xref>], to improve the efficiency of a CNN [<xref ref-type="bibr" rid="ref-65">65</xref>]. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates the difference between DSC and standard convolution. The parameter <inline-formula id="ieqn-17">
<mml:math id="mml-ieqn-17"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> of the standard convolution is calculated as follows:</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The difference between DSC and standard convolution</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-4.tif"/>
</fig>
<p><disp-formula id="eqn-3"><label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-18">
<mml:math id="mml-ieqn-18"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:math>
</inline-formula> is the number of input channels and <inline-formula id="ieqn-19">
<mml:math id="mml-ieqn-19"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</inline-formula> is the number of output channels of the structure, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4a</xref>.</p>
<p>As shown in <xref ref-type="fig" rid="fig-4">Fig. 4b</xref>, the parameter <inline-formula id="ieqn-20">
<mml:math id="mml-ieqn-20"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> of the depthwise convolution is calculated as follows:<disp-formula id="eqn-4"><label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:math>
</disp-formula></p>
<p>The ratio of standard convolution parameters to those of DSC is as follows:<disp-formula id="eqn-5"><label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>+</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>9</mml:mn></mml:mfrac></mml:mrow></mml:mstyle></mml:mstyle></mml:mstyle></mml:mstyle></mml:math>
</disp-formula></p>
<p>As shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, the parameter of DSC is 90% smaller compared with that of standard convolution. Therefore, the use of DSC instead of standard convolution in YOLOv3 can effectively improve efficiency.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Multi-Scale Feature Fusion by Dilated-Down Sampling (DDS)</title>
<p>The diagram of DDS is shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, which is the main component of the backbone network in MDN-YOLOv3. As depicted in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, a series of dilated convolution blocks are utilized in the network. There are three steps for DDS blocks to extract features. First, dilated convolution blocks with a stride of two and different dilation rates of one, two, and three are used to extract multi-scale features from the feature maps. Second, the dilated convolution block is combined using the concatenate operation to fuse features on different receptive scales. Finally, a 1 &#x00D7; 1 convolution layer is used to integrate every channel of the feature maps. The output of DDS contains features of various receptive fields to describe concrete cracks.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Block of the DDS</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Attention Module of CBAM</title>
<p>The attention module of the Convolutional Block Attention Module (CBAM) is introduced into the backbone network, and the structure of this module is shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. CBAM compresses feature maps using max-pooling (MaxPool) along with average-pooling (AvgPool) for channel and spatial attention.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The CBAM in bottleneck</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-6.tif"/>
</fig>
</sec>
<sec id="s3_1_4">
<label>3.1.4</label>
<title>Optimization of the Predictive Layer of the Backbone Module with a K-Means Cluster</title>
<p>The classification of the anchor box is a key parameter in YOLOv3. To verify the best configuration of the anchor box in concrete crack detection, the K-means clustering algorithm is used to cluster the dataset of the concrete crack. The dataset of the concrete crack in this study contains 6066 figures with seven categories, including no crack, hole, mesh crack, oblique crack, transverse crack, vertical crack, and irregular crack. Average IoU is an index used to reflect the effect of clustering between anchors and real target boxes. For the dataset of crack detection in this study, the relationship between the average IoU and the clustering number K is shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The clustering number K &#x003D; 4 is chosen because the average IoU has not changed significantly with the increase in clusters when K &#x003E; 4 and the point after the great slope can be regarded as the optimal point. The K-means clustering results of the training set are shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The position of the four hollow marks is the center of the corresponding cluster, including (47, 137), (128, 133), (136, 89), and (137, 43). Thus, the anchor boxes of the MDN-YOLOv3 method are (47, 137), (128, 133), (136, 89), and (137, 43) for the concrete crack detection, while those in YOLOv3 are (35, 137), (51, 137), (83, 135), (109, 135) (136, 78), (136, 132), (136, 100), (137, 36), and (137, 54). In addition, the predictive layers of MDN-YOLOv3 are 8-fold and 16-fold downsampling layers, while the predictive layers are 8, 16, and 32-fold in YOLOv3. The optimization of the anchor configuration simplifies the prediction layer of the backbone module from three scales to two scales, resulting in higher efficiency compared with M3-YOLOv3.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Clustering result of concrete-crack dataset</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-7.tif"/>
</fig><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Clustering results of nine ground-truth boxes. (a) YOLOv3 (b) SciYOLOv3</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Structure of MDN-YOLOv3</title>
<p>Combining the superiority of DSC and DDS and the attention mechanism of CBAM, the multi-dilated network (MDN) is proposed. MDN is used as the backbone network of YOLOv3, which can be called MDN-YOLOv3. Furthermore, the number of effective feature layers is streamlined from three to two. The whole structure of the MDN-YOLOv3 is shown in <xref ref-type="fig" rid="fig-9">Figs. 9</xref> and <xref ref-type="fig" rid="fig-10">10</xref>. The main structures can be divided into three modules: an MDN backbone, an FPN, and a prediction module.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>(a) The structure of MDN-YOLOv3. (b) The structure of the bottleneck</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-9.tif"/>
</fig><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Network details of MDN-YOLOv3</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-10.tif"/>
</fig>
<p>The MDN serves as the backbone network for extracting multi-scale features. Specifically, crosswise MDN structures are added between bottleneck modules to extract and fuse multi-scale features from input images. In addition, the CBAM attention mechanism is introduced to enhance feature extraction ability. The output matrix dimensions of the effective feature layers in the MDN backbone are 1/16 and 1/8 of the original image size, resulting in feature maps of 52 <inline-formula id="ieqn-21">
<mml:math id="mml-ieqn-21"><mml:mo>&#x00D7;</mml:mo></mml:math>
</inline-formula> 52 and 26 <inline-formula id="ieqn-22">
<mml:math id="mml-ieqn-22"><mml:mo>&#x00D7;</mml:mo></mml:math>
</inline-formula> 26 pixels being output. These output feature maps can be used for prediction (see <xref ref-type="fig" rid="fig-9">Fig. 9a</xref>). The resulting output data are a matrix with eleven rows and one column that includes four coordinate parameters, six classification parameters, and one confidence parameter.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Splicing Concrete-Crack Dataset (SCCD)</title>
<p>The existing crack datasets of concrete structures often focus on a single crack or several adjacent cracks. These crack patterns tend to be somewhat repetitive and are not suitable for identifying intricate patterns in real-world situations. To address this limitation, we develop a new dataset named the SCCD to capture more diverse and complex crack patterns. Overall, 6066 figures are collected and divided into seven categories: no crack, hole, mesh crack, oblique crack, transverse crack, vertical crack, and irregular crack, as shown in <xref ref-type="fig" rid="fig-11">Figs. 11a</xref>&#x2013;<xref ref-type="fig" rid="fig-11">11g</xref>. Moreover, every nine images are spliced as a new image to increase the complexity of crack patterns, as shown in <xref ref-type="fig" rid="fig-11">Fig. 11h</xref>. Various types of crack models can be generated randomly at different locations in splicing images with variable backgrounds. The new 674 general splicing images are further divided into two parts: 500 images are used as the training set and 174 images are used as the test set.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Concrete surface detect dataset. (a) no crack, (b) hole, (c) mesh crack, (d) oblique crack, (e) transverse crack, (f) vertical crack, (g) irregular crack, (h) splicing picture</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-11.tif"/>
</fig>
</sec>
<sec id="s5">
<label>5</label>
<title>Superiorities of MDN-YOLOv3</title>
<sec id="s5_1">
<label>5.1</label>
<title>Training Environment</title>
<p>This experiment is conducted on a Windows 10 system using PyTorch. The parameter settings for the training experiments on this platform are presented in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label>
<caption>
<title>Experimental platform configuration</title></caption>
<table><colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Attribute</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Operating system</td>
<td>Windows 10</td>
</tr>
<tr>
<td>CPU</td>
<td>Intel(R) Xeon(R) Gold 5222 CPU @ 3.80 GHz 3.79 GHz</td>
</tr>
<tr>
<td>GPU</td>
<td>NVIDIA Quadro P2200</td>
</tr>
<tr>
<td>RAM</td>
<td>64.0 GB</td>
</tr>
<tr>
<td rowspan="4">Programming environment</td>
<td>Anaconda3</td>
</tr>
<tr>
<td>CUDA10.2</td>
</tr>
<tr>
<td>Python3.6</td>
</tr>
<tr>
<td>PyTorch</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Training Parameters</title>
<p>When the MDN-YOLOv3 is trained, the learning rate is set to 0.001 and kept constant throughout the training process. The epoch, batch size, and weight decay numbers are set to 50, 4, and 0.0005, respectively.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Evaluate Indicators</title>
<p>The recall rate [<xref ref-type="bibr" rid="ref-66">66</xref>] and prediction rate [<xref ref-type="bibr" rid="ref-67">67</xref>] are calculated to evaluate the identification results of the MDN-YOLOv3. The definitions of these two parameters are<disp-formula id="eqn-6"><label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mrow></mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow></mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula><disp-formula id="eqn-7"><label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow></mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula></p>
<p>where <italic>TP</italic> represents the number of positive samples with a correct prediction, <italic>FP</italic> represents the number of wrong predictions, and <italic>FN</italic> represents the number of positive samples with a failed prediction.</p>
<p>In addition, two indexes, the mean average precision (<italic>mAP</italic>) at an IoU threshold of 0.5 and the frames per second (<italic>FPS</italic>), are calculated to quantitatively evaluate the calculation accuracy and speed of the network, respectively. The definition of <italic>mAP</italic> can be expressed as<disp-formula id="eqn-8"><label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mi>A</mml:mi><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mrow><mml:mi>N</mml:mi></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-9"><label>(9)</label>
<mml:math id="mml-eqn-9" display="block"><mml:mi>A</mml:mi><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mn>0</mml:mn><mml:mn>1</mml:mn></mml:msubsup><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math>
</disp-formula></p>
<p>where <italic>AP</italic><sub><italic>i</italic></sub> is the average accuracy of the i-th crack pattern. Generally, a higher <italic>AP</italic> value shows better classification results for a model.</p>
<p>The <inline-formula id="ieqn-23">
<mml:math id="mml-ieqn-23"><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>S</mml:mi></mml:math>
</inline-formula> represents the identification speed and can be expressed as</p>
<p><disp-formula id="eqn-10"><label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-24">
<mml:math id="mml-ieqn-24"><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math>
</inline-formula> is the number of images detected, and <inline-formula id="ieqn-25">
<mml:math id="mml-ieqn-25"><mml:mi>T</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:math>
</inline-formula> is the sum time of the detected images.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Comparison with Existing Typical Models</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> displays the structures utilized in every model. <xref ref-type="fig" rid="fig-12">Fig. 12</xref> illustrates the <italic>P&#x2013;R</italic> curves of six concrete cracks for various models, and their <italic>AP</italic> is calculated using <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref> and summarized in <xref ref-type="table" rid="table-3">Table 3</xref>. The highest <italic>AP</italic> among all of the compared models is observed for the hole, irregular crack, oblique crack, and transverse crack of MDN-YOLOv3; however, the mesh crack&#x2019;s <italic>AP</italic> is slightly lower than that of YOLOv3 and M3-YOLOv3, while the vertical crack&#x2019;s <italic>AP</italic> is a bit lower than that of YOLOv3, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. The <italic>mAP</italic> of the models is determined by calculating the respective <italic>APs</italic>, as depicted in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>. While MDN-YOLOv3 and YOLOv3 have the highest <italic>mAP</italic> values, M2-YOLOv3 has the lowest value among them all.</p>
<table-wrap id="table-2"><label>Table 2</label>
<caption>
<title>Comparison of model structures used in every model</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Structure</th>
<th>YOLOv3</th>
<th>M2-YOLOv3</th>
<th>M3-YOLOv3</th>
<th>MDN-YOLOv3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Backbone</td>
<td>Darknet-53</td>
<td>MobileNetv2</td>
<td>MobileNetv3</td>
<td>MDN</td>
</tr>
<tr>
<td>Bottleneck</td>
<td>N/A</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
</tr>
<tr>
<td>DSC</td>
<td>N/A</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
</tr>
<tr>
<td>Attention mechanism</td>
<td>N/A</td>
<td>N/A</td>
<td>&#x221A; (SE)</td>
<td>&#x221A; (CBAM)</td>
</tr>
<tr>
<td>DDS</td>
<td>N/A</td>
<td>N/A</td>
<td>N/A</td>
<td>&#x221A;</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-12">
<label>Figure 12</label>
<caption>
<title><italic>P&#x2013;R</italic> curves of six concrete cracks for different models: (a) YOLOv3, (b) M2-YOLOv3, (c) M3-YOLOv3, (d) MDN-YOLOv3</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-12.tif"/>
</fig><table-wrap id="table-3"><label>Table 3</label>
<caption>
<title><italic>AP</italic> for every crack type identified by the models</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Crack type</th>
<th colspan="4">AP (%)</th>
</tr>
<tr>
<th>YOLOv3</th>
<th>M2-YOLOv3</th>
<th>M3-YOLOv3</th>
<th>MDN-YOLOv3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Hole</td>
<td>79.02</td>
<td>80.03</td>
<td>79.47</td>
<td>84.24</td>
</tr>
<tr>
<td>Irregular crack</td>
<td>29.95</td>
<td>21.23</td>
<td>21.74</td>
<td>37.63</td>
</tr>
<tr>
<td>Mesh crack</td>
<td>59.59</td>
<td>40.54</td>
<td>48.14</td>
<td>42.67</td>
</tr>
<tr>
<td>Oblique crack</td>
<td>83.51</td>
<td>79.71</td>
<td>75.90</td>
<td>85.43</td>
</tr>
<tr>
<td>Transverse crack</td>
<td>88.43</td>
<td>72.88</td>
<td>87.05</td>
<td>89.65</td>
</tr>
<tr>
<td>Vertical crack</td>
<td>87.13</td>
<td>76.15</td>
<td>82.62</td>
<td>86.12</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title><italic>AP</italic> and <italic>mAP</italic> value comparison of different models</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-13.tif"/>
</fig>
<p><xref ref-type="table" rid="table-4">Table 4</xref> summarizes the evolution indicators of <italic>mAP</italic> and <italic>FPS</italic>. The MDN-YOLOv3 has a <italic>mAP</italic> of 70.96%, which is almost identical to YOLOv3&#x2019;s 71.27%. This value is significantly higher than that of M2-YOLOv3 and M3-YOLOv3. In addition, the <italic>FPS</italic> of MDN-YOLOv3 is 40.73 (f/s), making it faster by 41.5%, 8.25%, and 1.92% compared with YOLOv3, M2-YOLOv3, and M3-YOLOv3, respectively. Moreover, the model size of MDN-YOLOv3 is only 6.92 (MB), which makes it smaller by significant margins of 97.2%, 27.8%, and 57.7% when compared with the YOLOv3, M2-YOLOv3, and M3-YOLOv3 models, respectively.</p>
<table-wrap id="table-4"><label>Table 4</label>
<caption>
<title>Evaluate indicators of identification of the models</title></caption>
<table><colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Evaluate indicators</th>
<th>YOLOv3</th>
<th>M2-YOLOv3</th>
<th>M3-YOLOv3</th>
<th>MDN- YOLOv3</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>mAP</italic> (%)</td>
<td>71.27</td>
<td>61.76</td>
<td>65.82</td>
<td>70.96</td>
</tr>
<tr>
<td>FPS (f/s)</td>
<td>23.81</td>
<td>37.37</td>
<td>39.96</td>
<td>40.73</td>
</tr>
<tr>
<td>Parameter size (MB)</td>
<td>246.42</td>
<td>9.58</td>
<td>16.35</td>
<td>6.92</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To illustrate every model&#x2019;s performance vividly, a bubble diagram has been made in <xref ref-type="fig" rid="fig-14">Fig. 14</xref> where the bubble size represents every model&#x2019;s size, the abscissa shows the <italic>mAP</italic>, and the ordinate represents the corresponding <italic>FPS</italic>. As shown in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>, the bubble of MDN-YOLOv3 is the smallest and is located in the top right of the diagram. In other words, the MDN-YOLOv3 has the smallest model size and the fastest identification speed. Furthermore, the recognition accuracy of MDN-YOLOv3 is near that of M3-YOLOv3 and better than that of other comparison models.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Comparison of <italic>mAP</italic>, <italic>FPS</italic>, and model size</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-14.tif"/>
</fig>
<p>In addition, the cracks in the test set are detected and classified with corresponding confidence, as shown in <xref ref-type="fig" rid="fig-15">Fig. 15</xref>. The location and category of the cracks are clearly identified using the improved MDN-YOLOv3 method.</p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Crack detection results using MDN-YOLOv3</title></caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_42388-fig-15.tif"/>
</fig>
<p>In summary, the proposed MDN-YOLOv3 identifies cracks quickly while maintaining high accuracy levels despite its small model size.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions</title>
<p>The identification of crack locations and patterns is of great significance in evaluating the service performance of concrete structures. To address this issue, this study proposes an improved YOLOv3 algorithm for the rapid and accurate identification of the surface cracks of concrete structures. The proposed model is verified using a self-built training set. The main conclusions are presented as follows:</p>
<p>(1) The proposed MDN-YOLOv3 has lightweight network structures and robust feature extraction capability. The DSC and proposed DDS structure can effectively extract and fuse the features of cracks.</p>
<p>(2) The proposed MDN-YOLOv3 has a model size of 6.92 M and an identification speed of 40.73 FPS, which is 97.2% smaller in size and 41.5% faster in speed than YOLOv3. This finding indicates that the proposed MDN-YOLOv3 is more suitable for applications on mobile devices due to the faster calculation speed and smaller memory of these devices.</p>
<p>In the future, a better crack identification scheme should be studied to evaluate the applications of proposed models, such as temperature cracks, concrete shrinkage cracks, and external load cracks. In addition, concrete crack images from complex scenes such as underwater concrete and tunnel concrete require further research and development.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors are grateful for the International Science &#x0026; Technology Cooperation Project of Jiangsu Province (BZ2022010), the Jiangsu-Czech Bilateral Co-Funding R&#x0026;D Project (No. BZ2023011), and &#x201C;The Belt and Road&#x201D; Innovative Talents Exchange Foreign Experts Project (No. DL2023019001L).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Maosen Cao, Dragoslav Sumarac; data collection: Alfred Strauss, Yehui Shi; analysis and interpretation of results: Haoan Gu, Kai Zhu, Alfred Strauss; draft manuscript preparation: Haoan Gu, Maosen Cao. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the corresponding author upon reasonable request.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Srikanth</surname>, <given-names>I.</given-names></string-name>, <string-name><surname>Arockiasamy</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Deterioration models for prediction of remaining useful life of timber and concrete bridges: A review</article-title>. <source>Journal of Traffic and Transportation Engineering (English Edition)</source><italic>,</italic> <volume>7</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>152</fpage>&#x2013;<lpage>173</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtte.2019.09.005</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matsuoka</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Tokunaga</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Kaito</surname>, <given-names>K.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Bayesian estimation of instantaneous frequency reduction on cracked concrete railway bridges under high-speed train passage</article-title>. <source>Mechanical Systems and Signal Processing</source><italic>,</italic> <volume>161</volume><italic>(</italic><issue>9</issue><italic>),</italic> <fpage>107944</fpage>. <pub-id pub-id-type="doi">10.1016/j.ymssp.2021.107944</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Deng</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>H. L.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Mechanical property deterioration of the prefabricated concrete slab in mixed passenger and freight railway tracks</article-title>. <source>Construction and Building Materials</source><italic>,</italic> <volume>208</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>622</fpage>&#x2013;<lpage>637</lpage>. <pub-id pub-id-type="doi">10.1016/j.conbuildmat.2019.03.039</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dung</surname>, <given-names>C. V.</given-names></string-name>, <string-name><surname>Le</surname>, <given-names>D. A.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Autonomous concrete crack detection using deep fully convolutional neural network</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>99</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>52</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.11.028</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhong</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>M. Y.</given-names></string-name>, <string-name><surname>Zhai</surname>, <given-names>Y. Y.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Assessment of the feasibility of detecting concrete cracks in images acquired by unmanned aerial vehicles</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>89</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>49</fpage>&#x2013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.01.005</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Computer vision-based concrete crack detection using U-net fully convolutional networks</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>104</volume><italic>,</italic> <fpage>129</fpage>&#x2013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2019.04.005</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zerwer</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Polak</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Santamarina</surname>, <given-names>J. C.</given-names></string-name></person-group> (<year>2005</year>). <article-title>Detection of surface breaking cracks in concrete members using Rayleigh waves</article-title>. <source>Journal of Environmental &#x0026; Engineering Geophysics</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>295</fpage>&#x2013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.2113/JEEG10.3.295</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aghajanzadeh</surname>, <given-names>S. M.</given-names></string-name>, <string-name><surname>Mirzabozorg</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Concrete fracture process modeling by combination of extended finite element method and smeared crack approach</article-title>. <source>Theoretical and Applied Fracture Mechanics</source><italic>,</italic> <volume>101</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>306</fpage>&#x2013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1016/j.tafmec.2019.03.012</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdelkhalek</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Zayed</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Comprehensive inspection system for concrete bridge deck application: Current situation and future needs</article-title>. <source>Journal of Performance of Constructed Facilities</source><italic>,</italic> <volume>34</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>03120001</fpage>. <pub-id pub-id-type="doi">10.1061/(ASCE)CF.1943-5509.0001484</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2023</year>). <article-title>Automatic detection of concrete cracks from images using Adam-SqueezeNet deep learning model</article-title>. <source>Frattura ed Integrit&#x00E0; Strutturale</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>65</issue><italic>),</italic> <fpage>289</fpage>&#x2013;<lpage>299</lpage>. <pub-id pub-id-type="doi">10.3221/IGF-ESIS.65.19</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ohtsu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Shigeishi</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Iwase</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Koyanagit</surname>, <given-names>W.</given-names></string-name></person-group> (<year>1991</year>). <article-title>Determination of crack location, type and orientation ina concrete structures by acoustic emission</article-title>. <source>Magazine of Concrete Research</source><italic>,</italic> <volume>43</volume><italic>(</italic><issue>155</issue><italic>),</italic> <fpage>127</fpage>&#x2013;<lpage>134</lpage>. <pub-id pub-id-type="doi">10.1680/macr.1991.43.155.127</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goszczy&#x0144;ska</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>&#x015A;wit</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Tr&#x0105;mpczy&#x0144;ski</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Krampikowska</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Tworzewska</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2012</year>). <article-title>Experimental validation of concrete crack identification and location with acoustic emission method</article-title>. <source>Archives of Civil and Mechanical Engineering</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>23</fpage>&#x2013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1016/j.acme.2012.03.004</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Ahn</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Shin</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Sim</surname>, <given-names>S. H.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Crack and noncrack classification from concrete surface images using machine learning</article-title>. <source>Structural Health Monitoring</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>725</fpage>&#x2013;<lpage>738</lpage>. <pub-id pub-id-type="doi">10.1177/1475921718768747</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Lu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Lee</surname>, <given-names>V. C. S.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Concrete crack detection with handwriting script interferences using faster region-based convolutional neural network</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>35</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>373</fpage>&#x2013;<lpage>388</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12497</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Murao</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Nomura</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Furuta</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>C. W.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Concrete crack detection using UAV and deep learning</article-title>. <conf-name>Proceedings of the 13th International Conference on Applications of Statistics and Probability in Civil Engineering</conf-name>, <publisher-loc>Seoul, South Korea</publisher-loc>, <publisher-name>ICASP</publisher-name>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Farhidzadeh</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Dehghan-Niri</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Salamone</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Luna</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Whittaker</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2013</year>). <article-title>Monitoring crack propagation in reinforced concrete shear walls by acoustic emission</article-title>. <source>Journal of Structural Engineering</source><italic>,</italic> <volume>139</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>04013010</fpage>. <pub-id pub-id-type="doi">10.1061/(ASCE)ST.1943-541X.0000781</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>IAEA</collab></person-group> (<year>2022</year>). <source>Guidebook on non-destructive testing of concrete structures</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>International Atomic Energy Agency</publisher-name>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karthick</surname>, <given-names>S. P.</given-names></string-name>, <string-name><surname>Muralidharan</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Saraswathy</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Kwon</surname>, <given-names>S. J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Effect of different alkali salt additions on concrete durability property</article-title>. <source>Journal of Structural Integrity and Maintenance</source><italic>,</italic> <volume>1</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>35</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1080/24705314.2016.1153338</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Cho</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Spencer</surname> <suffix>Jr</suffix>, <given-names>B. F.</given-names></string-name>, <string-name><surname>Fan</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Automated assessment of cracks on concrete surfaces using adaptive digital image processing</article-title>. <source>Smart Structures and Systems</source><italic>,</italic> <volume>14</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>719</fpage>&#x2013;<lpage>741</lpage>. <pub-id pub-id-type="doi">10.12989/sss.2014.14.4.719</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jahanshahi</surname>, <given-names>M. R.</given-names></string-name>, <string-name><surname>Kelly</surname>, <given-names>J. S.</given-names></string-name>, <string-name><surname>Masri</surname>, <given-names>S. F.</given-names></string-name>, <string-name><surname>Sukhatme</surname>, <given-names>G. S.</given-names></string-name></person-group> (<year>2009</year>). <article-title>A survey and evaluation of promising approaches for automatic image-based defect detection of bridge structures</article-title>. <source>Structure and Infrastructure Engineering</source><italic>,</italic> <volume>5</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>455</fpage>&#x2013;<lpage>486</lpage>. <pub-id pub-id-type="doi">10.1080/15732470801945930</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Feng</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Feng</surname>, <given-names>M. Q.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Computer vision for SHM of civil infrastructure: From dynamic response measurement to damage detection&#x2013;A review</article-title>. <source>Engineering Structures</source><italic>,</italic> <volume>156</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>105</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.engstruct.2017.11.018</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koch</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Georgieva</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Kasireddy</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Akinci</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Fieguth</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2015</year>). <article-title>A review on computer vision based defect detection and condition assessment of concrete and asphalt civil infrastructure</article-title>. <source>Advanced Engineering Informatics</source><italic>,</italic> <volume>29</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>196</fpage>&#x2013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/j.aei.2015.01.008</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname>, <given-names>W. J.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>K. C.</given-names></string-name>, <string-name><surname>Braham</surname>, <given-names>A. F.</given-names></string-name>, <string-name><surname>Qiu</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Pavement crack width measurement based on Laplace&#x2019;s equation for continuity and unambiguity</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>33</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>110</fpage>&#x2013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12319</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname>, <given-names>S. N.</given-names></string-name>, <string-name><surname>Jang</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Han</surname>, <given-names>C. S.</given-names></string-name></person-group> (<year>2007</year>). <article-title>Auto inspection system using a mobile robot for detecting concrete cracks in a tunnel</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>255</fpage>&#x2013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2006.05.003</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>German</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Brilakis</surname>, <given-names>I.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Visual retrieval of concrete crack properties for automated post-earthquake structural safety evaluation</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>874</fpage>&#x2013;<lpage>883</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2011.03.004</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lattanzi</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Miller</surname>, <given-names>G. R.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Robust automated concrete damage detection algorithms for field applications</article-title>. <source>Journal of Computing in Civil Engineering</source><italic>,</italic> <volume>28</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>253</fpage>&#x2013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1061/(ASCE)CP.1943-5487.0000257</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yamaguchi</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Hashimoto</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Fast crack detection method for large-size concrete surface images using percolation-based image processing</article-title>. <source>Machine Vision and Applications</source><italic>,</italic> <volume>21</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>797</fpage>&#x2013;<lpage>809</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-009-0189-8</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Moon</surname>, <given-names>H. G.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>J. H.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Intelligent crack detecting algorithm on the concrete crack image using neural network</article-title>. <conf-name>Proceedings of the 28th ISARC</conf-name>, pp. <fpage>1461</fpage>&#x2013;<lpage>1467</lpage>. <publisher-loc>Seoul, South Korea</publisher-loc>. <pub-id pub-id-type="doi">10.22260/ISARC2011/0279</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Qi</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Automatic crack detection and classification method for subway tunnel safety monitoring</article-title>. <source>Sensors</source><italic>,</italic> <volume>14</volume><italic>(</italic><issue>10</issue><italic>),</italic> <fpage>19307</fpage>&#x2013;<lpage>19328</lpage>. <pub-id pub-id-type="doi">10.3390/s141019307</pub-id>; <pub-id pub-id-type="pmid">25325337</pub-id></mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prasanna</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Dana</surname>, <given-names>K. J.</given-names></string-name>, <string-name><surname>Gucunski</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Basily</surname>, <given-names>B. B.</given-names></string-name>, <string-name><surname>La</surname>, <given-names>H. M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2016</year>). <article-title>Automated crack detection on concrete bridges</article-title>. <source>IEEE Transactions on Automation Science and Engineering</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>591</fpage>&#x2013;<lpage>599</lpage>. <pub-id pub-id-type="doi">10.1109/TASE.2014.2354314</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Cui</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Qi</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Meng</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Automatic road crack detection using random structured forests</article-title>. <source>IEEE Transactions on Intelligent Transportation Systems</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>3434</fpage>&#x2013;<lpage>3445</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2016.2552248</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Du</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Ru</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Recognition and evaluation of bridge cracks with modified active contour model and greedy search-based support vector machine</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>78</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>51</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2017.01.019</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>Y. Z</given-names></string-name>, <string-name><surname>Nie</surname>, <given-names>Z. H</given-names></string-name>, <string-name><surname>Ma</surname>, <given-names>H. W</given-names></string-name></person-group> (<year>2017</year>). <article-title>Structural damage detection with automatic feature-extraction through deep learning</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>32</volume><italic>(</italic><issue>12</issue><italic>),</italic> <fpage>1025</fpage>&#x2013;<lpage>1046</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12313</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Song</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Deep learning-based roadway crack classification using laser-scanned range images: A comparative study on hyperparameter selection</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>114</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>103171</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2020.103171</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>K. C. P.</given-names></string-name>, <string-name><surname>Fei</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>C.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Automated pixel-level pavement crack detection on 3D asphalt surfaces with a recurrent neural network</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>34</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>213</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12409</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>36.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nhat-Duc</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Nguyen</surname>, <given-names>Q. L.</given-names></string-name>, <string-name><surname>Tran</surname>, <given-names>V. D.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Automatic recognition of asphalt pavement cracks using metaheuristic optimized edge detection algorithms and convolution neural network</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>94</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>203</fpage>&#x2013;<lpage>213</lpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2018.07.008</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>37.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gupta</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Goodman</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Patel</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Hosfelt</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Sajeev</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Creating xBD: A dataset for assessing building damage from satellite imagery</article-title>. <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</conf-name>, pp. <fpage>10</fpage>&#x2013;<lpage>17</lpage>. <publisher-loc>Los Angeles, CA, USA</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1911.09296</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>38.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname>, <given-names>J. Z.</given-names></string-name>, <string-name><surname>Lu</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Khaitan</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Zaytseva</surname>, <given-names>V.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Building damage detection in satellite imagery using convolutional neural networks</article-title>. <pub-id pub-id-type="doi">10.48550/arXiv.1910.06444</pub-id></mixed-citation></ref>
<ref id="ref-39"><label>39.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dong</surname>, <given-names>C. Z.</given-names></string-name>, <string-name><surname>Catbas</surname>, <given-names>F. N.</given-names></string-name></person-group> (<year>2021</year>). <article-title>A review of computer vision-based structural health monitoring at local and global levels</article-title>. <source>Structural Health Monitoring</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>692</fpage>&#x2013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1177/1475921720935585</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>40.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>J. W.</given-names></string-name>, <string-name><surname>Shao</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2023</year>). <article-title>Crack detection and quantification for concrete structures using UAV and transformer</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>152</volume><italic>,</italic> <fpage>104929</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2023.104929</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>41.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Shao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Shu</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Automatic pixel-level crack detection and evaluation of concrete structures using deep learning</article-title>. <source>Structural Control and Health Monitoring</source><italic>,</italic> <volume>29</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>e2981</fpage>. <pub-id pub-id-type="doi">10.1002/stc.2981</pub-id></mixed-citation></ref>
<ref id="ref-42"><label>42.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kang</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Benipal</surname>, <given-names>S. S.</given-names></string-name>, <string-name><surname>Gopal</surname>, <given-names>D. L.</given-names></string-name>, <string-name><surname>Shao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Cha</surname>, <given-names>Y. J.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Hybrid pixel-level concrete crack segmentation and quantification across complex backgrounds using deep learning</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>118</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>103291</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2020.103291</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>43.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Cha</surname>, <given-names>Y. J.</given-names></string-name>, <string-name><surname>Choi</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2017</year>). <chapter-title>Vision-based concrete crack detection using a convolutional neural network</chapter-title>. <source>Conference Proceedings of the Society for Experimental Mechanics Series</source>, pp. <fpage>71</fpage>&#x2013;<lpage>73</lpage>. <publisher-loc>Rhodes, Greece</publisher-loc>. <pub-id pub-id-type="doi">10.1007/978-3-319-54777-0_9</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>44.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fu</surname>, <given-names>R. H.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Z. J.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>M. S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Enhanced intelligent identification of concrete cracks using multi-layered image preprocessing-aided convolutional neural networks</article-title>. <source>Sensors</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>7</issue><italic>),</italic> <fpage>2021</fpage>. <pub-id pub-id-type="doi">10.3390/s20072021</pub-id>; <pub-id pub-id-type="pmid">32260302</pub-id></mixed-citation></ref>
<ref id="ref-45"><label>45.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Ding</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Lin</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Duan</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Continual-learning-based framework for structural damage recognition</article-title>. <source>Structural Control and Health Monitoring</source><italic>,</italic> <volume>29</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>e3093</fpage>. <pub-id pub-id-type="doi">10.1002/stc.3093</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>46.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Wei</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Tang</surname>, <given-names>Z.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>The state of the art of data science and engineering in structural health monitoring</article-title>. <source>Engineering</source><italic>,</italic> <volume>5</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>234</fpage>&#x2013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1016/j.eng.2018.11.027</pub-id></mixed-citation></ref>
<ref id="ref-47"><label>47.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Fast R-CNN</article-title>. <conf-name>Proceedings of the IEEE International Conference on Vision</conf-name>, pp. <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>. <publisher-loc>Santiago, Chile</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1504.08083</pub-id></mixed-citation></ref>
<ref id="ref-48"><label>48.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname>, <given-names>S. Q.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>K. M.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Faster R-CNN: Towards real-time object detection with region proposal networks</article-title>. <source>Advances in Neural Information Processing Systems</source><italic>,</italic> <volume>28</volume>. <pub-id pub-id-type="doi">10.48550/arXiv.1506.01497</pub-id></mixed-citation></ref>
<ref id="ref-49"><label>49.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Guo</surname>, <given-names>P. W.</given-names></string-name>, <string-name><surname>Meng</surname>, <given-names>W. N.</given-names></string-name>, <string-name><surname>Bao</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Automatic identification and quantification of dense microcracks in high-performance fiber-reinforced cementitious composites through deep learning-based computer vision</article-title>. <source>Cement and Concrete Research</source><italic>,</italic> <volume>148</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>106532</fpage>. <pub-id pub-id-type="doi">10.1016/j.cemconres.2021.106532</pub-id></mixed-citation></ref>
<ref id="ref-50"><label>50.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Redmon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Divvala</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Girshick</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Farhadi</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2016</year>). <article-title>You only look once: Unified, real-time object detection</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>779</fpage>&#x2013;<lpage>788</lpage>. <publisher-loc>Las Vegas, USA</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1506.02640</pub-id></mixed-citation></ref>
<ref id="ref-51"><label>51.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Park</surname>, <given-names>S. E.</given-names></string-name>, <string-name><surname>Eem</surname>, <given-names>S. H.</given-names></string-name>, <string-name><surname>Jeon</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Concrete crack detection and quantification using deep learning and structured light</article-title>. <source>Construction and Building Materials</source><italic>,</italic> <volume>252</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>119096</fpage>. <pub-id pub-id-type="doi">10.1016/j.conbuildmat.2020.119096</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>52.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nie</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Pavement crack detection based on yolo v3</article-title>. <conf-name>2nd International Conference on Safety Produce Informatization (IICSPI)</conf-name>, pp. <fpage>327</fpage>&#x2013;<lpage>330</lpage>. <publisher-loc>Chongqing, China</publisher-loc>, <publisher-name>IEEE</publisher-name>. <pub-id pub-id-type="doi">10.1109/IICSPI48186.2019.9095956</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>53.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Gu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2023</year>). <article-title>Automatic recognition of pavement cracks from combined GPR B-scan and C-scan images using multiscale feature fusion deep neural networks</article-title>. <source>Automation in Construction</source><italic>,</italic> <volume>146</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>104698</fpage>. <pub-id pub-id-type="doi">10.1016/j.autcon.2022.104698</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>54.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Gu</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Y.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2022</year>). <article-title>Novel YOLOv3 model with structure and hyperparameter optimization for detection of pavement concealed cracks in GPR images</article-title>. <source>IEEE Transactions on Intelligent Transportation Systems</source><italic>,</italic> <volume>23</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>22258</fpage>&#x2013;<lpage>22268</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2022.3174626</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>55.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Crack detection of track plate based on YOLO</article-title>. <conf-name>2019 12th International Symposium on Computational Intelligence and Design (ISCID)</conf-name>, vol. <volume>2</volume>, pp. <fpage>15</fpage>&#x2013;<lpage>18</lpage>. <publisher-loc>Hangzhou, China</publisher-loc>, <publisher-name>IEEE</publisher-name>. <pub-id pub-id-type="doi">10.1109/ISCID.2019.10086</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>56.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Han</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>YOLOv3-Lite: A lightweight crack detection network for aircraft structure based on depthwise separable convolutions</article-title>. <source>Applied Sciences</source><italic>,</italic> <volume>9</volume><italic>(</italic><issue>18</issue><italic>),</italic> <fpage>3781</fpage>. <pub-id pub-id-type="doi">10.3390/app9183781</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>57.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Early recognition of tomato gray leaf spot disease based on MobileNetv2-YOLOv3 model</article-title>. <source>Plant Methods</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>83</fpage>. <pub-id pub-id-type="doi">10.1186/s13007-020-00624-2</pub-id>; <pub-id pub-id-type="pmid">32523613</pub-id></mixed-citation></ref>
<ref id="ref-58"><label>58.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Kang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Feng</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Automatic recognition of dairy cow mastitis from thermal images by a deep learning detector</article-title>. <source>Computers and Electronics in Agriculture</source><italic>,</italic> <volume>178</volume><italic>,</italic> <fpage>105754</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2020.105754</pub-id></mixed-citation></ref>
<ref id="ref-59"><label>59.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Howard</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Sandler</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Chu</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>L. C.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Searching for MobileNetV3</article-title>. <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, pp. <fpage>1314</fpage>&#x2013;<lpage>1324</lpage>. <publisher-loc>Seoul, South Korea</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1905.02244</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>60.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sandler</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Howard</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Zhu</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Zhmoginov</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>L. C.</given-names></string-name></person-group> (<year>2018</year>). <article-title>MobileNetv2: Inverted residuals and linear bottlenecks</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>4510</fpage>&#x2013;<lpage>4520</lpage>. <publisher-loc>Salt Lake City, USA</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1801.04381</pub-id></mixed-citation></ref>
<ref id="ref-61"><label>61.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Redmon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Farhadi</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Yolov3: An incremental improvement</article-title>. <pub-id pub-id-type="doi">10.48550/arXiv.1804.02767</pub-id></mixed-citation></ref>
<ref id="ref-62"><label>62.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Redmon</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Farhadi</surname>, <given-names>A.</given-names></string-name></person-group> (<year>2017</year>). <article-title>YOLO9000: Better, faster, stronger</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>7263</fpage>&#x2013;<lpage>7271</lpage>. <publisher-loc>Hawaii, USA</publisher-loc>. <pub-id pub-id-type="doi">10.48550/arXiv.1612.08242</pub-id></mixed-citation></ref>
<ref id="ref-63"><label>63.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cui</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Dai</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Intelligent recognition of erosion damage to concrete based on improved YOLO-v3</article-title>. <source>Materials Letters</source><italic>,</italic> <volume>302</volume><italic>(</italic><issue>3&#x2013;4</issue><italic>),</italic> <fpage>130363</fpage>. <pub-id pub-id-type="doi">10.1016/j.matlet.2021.130363</pub-id></mixed-citation></ref>
<ref id="ref-64"><label>64.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Cai</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2020</year>). <article-title>On bridge surface crack detection based on an improved YOLO v3 algorithm</article-title>. <source>IFAC-PapersOnLine</source><italic>,</italic> <volume>53</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>8205</fpage>&#x2013;<lpage>8210</lpage>. <pub-id pub-id-type="doi">10.1016/j.ifacol.2020.12.1994</pub-id></mixed-citation></ref>
<ref id="ref-65"><label>65.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Gong</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Automatic detection method of tunnel lining multi-defects via an enhanced you only look once network</article-title>. <source>Computer-Aided Civil and Infrastructure Engineering</source><italic>,</italic> <volume>37</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>762</fpage>&#x2013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12836</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>66.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bu&#x010D;ko</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Lieskovsk&#x00E1;</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Z&#x00E1;bovsk&#x00E1;</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Z&#x00E1;bovsk&#x00FD;</surname>, <given-names>M.</given-names></string-name></person-group> (<year>2022</year>). <article-title>Computer vision based pothole detection under challenging conditions</article-title>. <source>Sensors</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>22</issue><italic>),</italic> <fpage>8878</fpage>. <pub-id pub-id-type="doi">10.3390/s22228878</pub-id>; <pub-id pub-id-type="pmid">36433474</pub-id></mixed-citation></ref>
<ref id="ref-67"><label>67.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Luo</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Qiu</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Real-time pattern-recognition of GPR images with YOLO v3 implemented by tensorflow</article-title>. <source>Sensors</source><italic>,</italic> <volume>20</volume><italic>(</italic><issue>22</issue><italic>),</italic> <fpage>6476</fpage>. <pub-id pub-id-type="doi">10.3390/s20226476</pub-id>; <pub-id pub-id-type="pmid">33198420</pub-id></mixed-citation></ref>
</ref-list>
</back>
</article>


















