<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">42224</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.042224</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>C2Net-YOLOv5: A Bidirectional Res2Net-Based Traffic Sign Detection Algorithm</article-title>
<alt-title alt-title-type="left-running-head">C2Net-YOLOv5: A Bidirectional Res2Net-Based Traffic Sign Detection Algorithm</alt-title>
<alt-title alt-title-type="right-running-head">C2Net-YOLOv5: A Bidirectional Res2Net-Based Traffic Sign Detection Algorithm</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Xiujuan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Tian</surname><given-names>Yiqi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>tianyiqi@emails.bjut.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Zheng</surname><given-names>Kangfeng</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Chutong</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Faculty of Information Technology, Beijing University of Technology</institution>, <addr-line>Beijing, 100124</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Cyberspace Security, Beijing University of Posts and Telecommunications</institution>, <addr-line>Beijing, 100048</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Fan Gongxiu Honors College, Beijing University of Technology</institution>, <addr-line>Beijing, 100124</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yiqi Tian. Email: <email>tianyiqi@emails.bjut.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>29</day><month>11</month><year>2023</year></pub-date>
<volume>77</volume>
<issue>2</issue>
<fpage>1949</fpage>
<lpage>1965</lpage>
<history>
<date date-type="received"><day>23</day><month>5</month><year>2023</year></date>
<date date-type="accepted"><day>28</day><month>9</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Wang et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Wang et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_42224.pdf"></self-uri>
<abstract>
<p>Rapid advancement of intelligent transportation systems (ITS) and autonomous driving (AD) have shown the importance of accurate and efficient detection of traffic signs. However, certain drawbacks, such as balancing accuracy and real-time performance, hinder the deployment of traffic sign detection algorithms in ITS and AD domains. In this study, a novel traffic sign detection algorithm was proposed based on the bidirectional Res2Net architecture to achieve an improved balance between accuracy and speed. An enhanced backbone network module, called C2Net, which uses an upgraded bidirectional Res2Net, was introduced to mitigate information loss in the feature extraction process and to achieve information complementarity. Furthermore, a squeeze-and-excitation attention mechanism was incorporated within the channel attention of the architecture to perform channel-level feature correction on the input feature map, which effectively retains valuable features while removing non-essential features. A series of ablation experiments were conducted to validate the efficacy of the proposed methodology. The performance was evaluated using two distinct datasets: the Tsinghua-Tencent 100K and the CSUST Chinese traffic sign detection benchmark 2021. On the TT100K dataset, the method achieves precision, recall, and Map0.5 scores of 83.3&#x0025;, 79.3&#x0025;, and 84.2&#x0025;, respectively. Similarly, on the CCTSDB 2021 dataset, the method achieves precision, recall, and Map0.5 scores of 91.49&#x0025;, 73.79&#x0025;, and 81.03&#x0025;, respectively. Experimental results revealed that the proposed method had superior performance compared to conventional models, which includes the faster region-based convolutional neural network, single shot multibox detector, and you only look once version 5.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Target detection</kwd>
<kwd>traffic sign detection</kwd>
<kwd>autonomous driving</kwd>
<kwd>YOLOv5</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Key R&#x0026;D Program of China</funding-source>
<award-id>2017YFB0802803</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Beijing Natural Science Foundation</funding-source>
<award-id>4202002</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Traffic signs serve as important instructions and warnings on roads, guiding drivers to adhere to traffic rules and prevent accidents. The precise detection and identification of these signs are essential in assisted and autonomous driving systems. However, traffic sign detection faces two main challenges: first, the existing traffic sign detection algorithms usually require a substantial number of parameters and operations to achieve satisfactory results, making real-time performance unattainable; second, traffic sign detection algorithms encounter difficulties extracting sufficient features from the model. To address these challenges, this study proposes the C2Net-YOLOv5 model.</p>
<p>The you only look once version 5 (YOLOv5) model stands out as a lightweight, high-performance object detection framework. It combines the advantages of single-stage detection with improved network architecture and can detect objects in images in real-time scenes quickly and accurately. To maintain a comparable detection speed as the one-stage detector, the model uses the one-stage detector YOLOv5 as the basic model architecture. Furthermore, a bidirectional Res2Net [<xref ref-type="bibr" rid="ref-1">1</xref>] module is incorporated into the model&#x2019;s backbone for feature extraction, intended to solve the problem of inadequate extraction of features.</p>
<p>In this study, the designs are inspired by the human brain. While the convolutional neural networks (CNNs) are inspired by the human brain vision system, they do not precisely simulate the operational mode of the human brain. The human visual system entails complex multi-level processing within the cerebral cortex, hierarchical extraction of different features, and fine-grained information processing. Although CNNs draw on some characteristics of the human visual system, it is still a highly simplified and abstract model.</p>
<p>Similarly, the designs draw insights from features of the human brain system. The human brain exhibits a degree of symmetry, with two cerebral hemispheres on the left and right sides. These hemispheres serve different roles in processing visual information: the left brain is more inclined towards logical, analytical, and sequential processing, while the right brain specializes in spatial perception and positioning. Although the brain&#x2019;s hemispheres are functionally interconnected, they communicate and integrate information through the corpus callosum to ensure the implementation of comprehensive cognition and behavior. Therefore, the design adopts a symmetrical Res2Net structure to jointly integrate different information from both sides, complemented by attention mechanisms to filter features.</p>
<p>Unlike the existing YOLOv5 model, this model has been improved on the current version to mitigate the impact of scale invariance. Features of traffic signs within images are enhanced and extraneous background information is suppressed, rendering the model more resilient to the environment.</p>
<p>The method is tested using the Tsinghua-Tencent 100K (TT100K) and the CSUST Chinese traffic sign detection benchmark (CCTSDB) 2021 datasets, and the results demonstrated the method&#x2019;s validity. On the TT100K dataset, the method achieves precision, recall, and Map0.5 scores of 83.3&#x0025;, 79.3&#x0025;, and 84.2&#x0025;, respectively. Similarly, on the CCTSDB 2021 dataset, the method achieves precision, recall, and Map0.5 scores of 91.49&#x0025;, 73.79&#x0025;, and 81.03&#x0025;, respectively.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>This section presents a summary of existing research on traffic sign detection.</p>
<sec id="s2_1"><label>2.1</label><title>Traditional Traffic Sign Detection</title>
<p>Traditional traffic sign detection involves three key steps: region selection, feature extraction, and classification regression [<xref ref-type="bibr" rid="ref-2">2</xref>]. Li et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] introduced a method for road recognition by combining color-invariant-based image segmentation. Maldonado-Basc&#x00F3;n et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] presented an automatic road-sign detection and recognition system based on support vector machines (SVMs) which is able to detect and recognize circular, rectangular, triangular, and octagonal signs and, hence, covers all existing Spanish traffic-sign shapes. However, traditional target detection algorithms face two main challenges. First, the sliding window-based region selection method lacks precision, consumes excessive time, and the window is redundant. Second, manually extracted features are unstable in a dynamically changing environment.</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>Traffic Sign Detection Based on Deep Learning</title>
<p>Recently, the application of deep learning in traffic sign detection has gained significant traction. Notably, there are two distinctive methods: the two-stage and one-level traffic sign detection methods.</p>
<p>The two-stage traffic sign detection method is recognized for its high accuracy but relatively slower processing speed. In 2014, Girshick et al. introduced the region-based CNN (R-CNN) [<xref ref-type="bibr" rid="ref-5">5</xref>], a pioneering success in applying the deep learning method to object recognition. Building on this foundation, Fast R-CNN [<xref ref-type="bibr" rid="ref-6">6</xref>] was proposed in the subsequent year, aimed at enhancing the speed of R-CNN by unifying category judgment and frame regression through CNN implementation without requiring additional storage features. In 2017, Ren et al. proposed the Faster R-CNN [<xref ref-type="bibr" rid="ref-7">7</xref>], greatly improving the comprehensive performance. Additionally, Lin et al. proposed the feature pyramid network (FPN) [<xref ref-type="bibr" rid="ref-8">8</xref>], which uses feature maps of different resolutions to comprehend targets of different sizes. This method combined output features with shallow visual and deep-level semantic features through continuous upsampling and cross-layer fusion information. In 2018, Liu et al. proposed the path aggregation network (PAN) [<xref ref-type="bibr" rid="ref-9">9</xref>]; the original FPN is a one-way fusion from deep to shallow, but PAN is a bidirectional fusion from deep to shallow and <italic>vice versa</italic>. Additionally, Cai et al. introduced the cascade R-CNN method [<xref ref-type="bibr" rid="ref-10">10</xref>]. In 2019, Han et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] proposed a real-time small traffic sign detection approach based on a revised Faster R-CNN, which uses a small region proposal generator to extract the characteristics of small traffic signs and combine the revised architecture of Faster R-CNN with online hard examples mining (OHEM) to make the system more robust to locate the region of small traffic signs.</p>
<p>However, the one-level traffic sign detection method is faster and can achieve satisfactory accuracy. Notable algorithms in this category include the YOLO series proposed by Redmon et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] and the single shot multiBox detector (SSD) proposed by Liu et al. [<xref ref-type="bibr" rid="ref-13">13</xref>]. In 2018, Redmon et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed the YOLOv3, which integrates the ideas of current excellent detection frameworks, such as residual networks and feature fusion. You et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] proposed an end-to-end deep learning model for identifying traffic signs in high-definition pictures, which contains fundamental feature extraction and multitask learning. Kong et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] collected traffic signs in South Korea. They proposed a lightweight traffic sign detection method using cascaded CNN [<xref ref-type="bibr" rid="ref-17">17</xref>], which is hardware-friendly and reduces the computational complexity compared with Agone&#x2019;s proposed YOLOv2-tiny [<xref ref-type="bibr" rid="ref-18">18</xref>]. Yen et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed a CNN configured with area masks to resolve the occlusion problem in traffic sign detection. The method was highly effective in alleviating the occlusion problem. Siniosoglou et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed an auto-encoder model, which showed high precision in recognizing fuzzy traffic signs. Franzen et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] used a neural network trained in the frequency domain to detect traffic signs, greatly reducing the number of neurons compared to the traditional neural network. In 2020, Jocher et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed the YOLOv5, a progressive addition to the YOLO family of algorithms. Until now, YOLOv5 continues to undergo upgrades and iterations.</p>
<p>In 2021, Nagrath et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] designed the SSDMNV2 approach; it uses SSD as a detector and MobileNetV2 architecture as a framework for the classifier. This lightweight setup is suitable for real-time mask detection, even on embedded devices. Additionally, Pooja et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed a detection method that uses TensorFlow and OpenCV. At the same year, Du et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] established a target object grab setting model with the multi-target object and the anchor frame generation measurement strategy overcoming external environmental interference factors such as mutual interference between objects and changes in illumination. In 2022, Liu et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed a symmetrical traffic sign detection algorithm M-YOLO, for complex scenes. The algorithm optimizes the delay by reducing network computational overhead and speeding up feature extraction. Similarly, Loey et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a detection model consisting of two components. The first component is designed for feature extraction using ResNet50, while the second component is designed for the classification process of face masks using decision trees, SVMs, and ensemble algorithm. Yun et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a method of cluttering pose detection based on convolutional neural network with multiple self-powered sensors information. In 2023, Shi et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] proposed the cross-stage attention network module to enhance the feature extraction capability of the network. They used a dense neck structure for the comprehensive fusion of detail and semantic information. Liu et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposed a new key point assumption strategy based on the basis of the PvNet model. Meanwhile, a fusion method of pixel-by-pixel key point voting and depth image is applied to improve the performance of the model.</p>
<p>However, this study proposes the C2Net-YOLOv5 model for traffic sign detection based on the YOLOv5 framework, offering enhanced speed and suitability in real-time applications. The YOLO framework is developing rapidly; many have already found practical applications across various domains.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Algorithms</title>
<p>The YOLOv5 method is the best one-stage target detection method, distinguished by its computational efficiency and fast processing speed. In this study, the C2Net-YOLOv5 model was constructed based on the YOLOv5 framework, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>C2Net-YOLOv5 model</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-1.tif"/></fig>
<p>In the Backbone segment of the C2Net-YOLOv5, the main features are extracted through focus, convolution (Conv), C2Net, and spatial pyramid pooling. Focus is a special convolutional structure designed by YOLOv5&#x2019;s authors for multi-scale feature extraction in small target detection. C2Net enhances the network&#x2019;s ability to handle targets of varying scales, thereby improving detection performance. Conv refers to the convolution layer responsible for constructing the backbone network, feature extraction module, and classifier. It extracts features from the input feature graph and facilitates information processing and transformation.</p>
<p>The Neck segment adopts the combination of FPN and PAN concepts. The FPN fuses features from different scales and then performs prediction on the fused feature map. FPN is a feature pyramid network structure for target detection and semantic segmentation tasks. It enables the construction of a feature pyramid through a top-down feature propagation process and lateral connection. This mitigates target scale changes and facilitates small target detection. However, the PAN combines feature pyramids and path aggregation for comprehensive multi-scale information fusion. It leverages the backbone network to construct a multi-layer feature pyramid on the feature graph extracted at different stages of the backbone network. PAN efficiently captures both semantic and detailed characteristics of targets through effective multi-scale information. The combination of FPN and PAN uses the advantages of the two network structures to achieve more effective and robust feature representation and information fusion.</p>
<p>The Head module serves as the output layer, further extracting network features and transforming them into target detection boxes and category predictions for input images.</p>
<sec id="s3_1"><label>3.1</label><title>Construction of the C2Net Module</title>
<p>The C2Net module consists of three ConvBNSiLU modules, multiple Bottle2neck modules, and a Concat module, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The ConvBNSiLU module refers to the structural combination of Conv, batch normalization (BN), and sigmoid-weighted linear unit (SiLU). The Bottle2neck module will be introduced in <xref ref-type="sec" rid="s3_2">Section 3.2</xref>. The Concat module is a commonly used module in deep learning to connect multiple input features along specific dimensions.</p>
<p>In the earlier version of YOLOv5, the backbone network used the BottleneckCSP module for feature extraction. C2Net differs from the BottleneckCSP module; it removes the Conv module after the remaining output and replaces the activation function in the standard convolutional module with SiLU. The C2Net module is structured into two branches: one uses the specified multi-Bottle2neck stacks and ConvBNSiLU modules, while the other traverses a single ConvBNSiLU module. Subsequently, the branches were subjected to the Concat and ConvBNSiLU operations.</p>
<p>The advantages of the C2Net module are as follows: firstly, concatenating multiple convolutional layers boosts the model depth and width. This enhances the model&#x2019;s expressive ability to learn complex feature representations; secondly, the module uses convolution operations at varying scales to fuse multiple feature maps. This fusion of multi-scale features improves the model&#x2019;s perception of different target sizes and details. Additionally, the C2Net module leverages parallel computing through concatenating and convolving feature maps of different sizes, which optimizes computational efficiency. In summary, the C2Net module excels in multi-scale feature fusion and computational efficiency.</p>
</sec>
<sec id="s3_2"><label>3.2</label><title>Bottle2neck Module</title>
<p>To address the challenges of insufficient feature extraction in the model, a bidirectional Res2Net module is introduced into the Bottle2neck module, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2c</xref>. This module facilitates multi-scale feature extraction in two opposite directions, which enhances feature representation. This approach also rectifies the singularity orientation inherent in the Res2Net model. In this study, the addition of squeeze-and-excitation (SE) channel attention resolves the issues arising from varying channel importance during convolution pooling.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Res2Net, reverse Res2Net and Bottle2neck modules</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-2.tif"/></fig>
<p>The Bottle2neck module takes full advantage of Res2Net for multi-scale feature extraction. The standard 1-3-1 CNN layout is replaced with multi-scale residual layering architecture. This alteration shifts the main convolution in the middle from a single branch to a multi-branch configuration. By increasing the receptive fields within the block, different scale levels are captured layer by layer at a finer granularity, enhancing the CNN&#x2019;s ability to detect objects within images. In this study, the bidirectional Res2Net module was further used to conduct multi-scale feature extraction in two opposite directions to rectify the singularity in the direction of the Res2Net model.</p>
<sec id="s3_2_1"><label>3.2.1</label><title>Res2Net Module and Reverse Res2Net Module</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2a</xref> shows the Res2Net module used in this study. It constructs residual blocks of multiple branches within a single residual block. This module refines multi-scale features at a more granular level and expands the network&#x2019;s perceptual field. The steps are outlined as follows:
<list list-type="simple">
<list-item><label>1)</label><p>First, the module introduced a novel parameter, referred to as scale (denoted as s), which signifies the number of groups into which the feature map is divided.</p></list-item>
<list-item><label>2)</label><p>Next, the output features of the first 1&#x2009;&#x00D7;&#x2009;1 convolutional layer were divided by Res2Net into s equal groups based on channels, with each group having w channels, i.e., n&#x2009;&#x003D;&#x2009;s&#x2009;&#x00D7;&#x2009;w.</p></list-item>
<list-item><label>3)</label><p>Next, the second layer convolution kernel in the original Bottleneck block was divided by Res2Net into s groups, with each group having output channels w (similar to step 2). The convolution operation for each group is denoted <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>().</p></list-item>
<list-item><label>4)</label><p>For each group of features, after grouping <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, all groups corresponded to the convolution operation except the last group, which omitted the convolution operation <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>(), where i &#x2208; &#x007B;2,&#x2026;, s&#x007D;. Note <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the output of the convolution operation <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>(), then from the second group onwards, each convolution operation <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>() is preceded by the output of the preceding group <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and added to the features of the current group <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> forming a residual concatenation through <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mtext>K</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>(). This process continued until the penultimate set of features and can be expressed by the following equation:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mtext>&#x00A0;</mml:mtext><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mspace width="2em" /><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p></list-item>
<list-item><label>5)</label><p>Finally, the outputs corresponding to each group were channel concatenated and fed into a final layer of 1&#x2009;&#x00D7;&#x2009;1 convolutional layers to fuse the multi-scale features and obtain the block&#x2019;s output. This module&#x2019;s distinctive structure, characterized by a residual-like concatenation mechanism, is referred to as Res2Net.</p></list-item>
</list></p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2b</xref> shows the reverse Res2Net module, which is symmetrically aligned with the Res2Net module. This module aims to enhance feature representation and achieve complementary information.</p>
</sec>
<sec id="s3_2_2"><label>3.2.2</label><title>SE Channel Attention Module</title>
<p>The squeeze-and-excitation network (SENet) focuses on inter-channel relationships, enabling the model to automatically learn the importance of different channel features. SENet proposed the SE module as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The operational process is as follows:
<list list-type="simple">
<list-item><label>1)</label><p>Squeeze: Using global average pooling, the two-dimensional features (H&#x2009;&#x002A;&#x2009;W) of each channel were compressed into a real number. The resulting feature map transformed from [h, w, c]&#x003D;&#x003D;&gt; [1, 1, c], yielding channel-level global features.</p></list-item>
<list-item><label>2)</label><p>Excitation: Weight values were generated for each feature channel, establishing correlations between channels through two fully connected layers. The number of output weight values corresponds to the number of channels in the output feature map. The operation transformed from [1, 1, c] &#x003D;&#x003D;&gt; [1, 1, c], learned the relationships between different channels, and obtained the weights of different channels.</p></list-item>
<list-item><label>3)</label><p>Scale: The normalized weights obtained earlier were applied to each channel&#x2019;s features. This was achieved through channel-wise multiplication of weight coefficients, i.e., [h, w, c]&#x2009;&#x002A;&#x2009;[1, 1, c] &#x003D;&#x003D;&gt; [h, w, c].</p></list-item>
</list></p>
<fig id="fig-3"><label>Figure 3</label><caption><title>SE channel attention</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-3.tif"/></fig>
<p>Essentially, the SE module performs attention or gating operations within the channel dimension, which allows the model to focus more on channel features with higher information content while suppressing less important ones.</p>
<p>In <xref ref-type="fig" rid="fig-2">Fig. 2c</xref>, the Bottle2neck module uses the channel attention module. Given that channel weights for different channels often differ within an image, capturing this information enhances the model&#x2019;s overall information capacity and accuracy.</p>
</sec>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Experiments and Analysis of Results</title>
<p>This section presents the experimental evaluation of the proposed method. The proposed traffic sign detection method is implemented using the PyTorch deep learning framework. The method&#x2019;s efficacy was verified through testing on the TT100K and CCTSDB 2021 datasets.</p>
<sec id="s4_1"><label>4.1</label><title>Experimental Data</title>
<p>This study used the TT100K and CCTSDB 2021 [<xref ref-type="bibr" rid="ref-31">31</xref>] datasets. The TT100K is a traffic sign dataset produced by Tencent and Tsinghua, with a total of 100,000 images, of which 10,000 contains traffic signs. The training set consists of 6,150 images, while the test set consists of 3,071 images. Notably, the dataset exhibited category imbalance, as certain signs (e.g., landslides and villages ahead) were not visible in the city center. Additionally, there were 70 missing categories, i.e., 70 categories lacked instances, emphasizing the importance of meticulous data processing. Only categories containing more than 100 images were retained.</p>
<p>The CCTSDB 2021 dataset is a novel Chinese traffic sign detection benchmark proposed by the authors, which adds over 4,000 real traffic sign images and homologous detailed labels to CCTSDB 2017. Furthermore, it replaces many original easy-to-detect images with difficult samples to fit the complex and variable detection environment.</p>
</sec>
<sec id="s4_2"><label>4.2</label><title>Assessment Indicators</title>
<p>In evaluating the experimental results, box_loss, cls_loss, obj_loss, precision, recall, Map0.5, and F1-score were used as indicators for evaluating the proposed methods.</p>
<p>The cls_loss was used as the classification loss function. The model generated three prediction boxes for each N&#x2009;&#x002A;&#x2009;N grid cell containing nc classification probabilities. The formula for calculating cls_loss, as presented in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>, involves several components, such as label (representing values in the unique heat code label), &#x03B1; (the smoothing coefficient with a value ranging from 0 to 1), and nc (representing the total number of categories). The label probability matrix is denoted as matrix <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, while the prediction probability is represented as matrix P. The Binary CrossEntropy Loss (BCE Loss) for each numerical entry in the matrix is referred to as <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In this context, nc represents the dataset category, N represents the grid size, z represents the z-th anchor in the grid, x represents the abscissa position, y represents the ordinate position, and t represents the t-th category.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">label</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mrow><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">smooth</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">false</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">false</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:munder><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The box_loss was used as the localization loss to indicate the deviation between the prediction and calibration boxes. The formula for calculating box_loss, as presented in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, consists of a constant <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">coord</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> that balances the loss of position and size. The variables <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and B represents the number of grid cells and candidate boxs, respectively. Furthermore, x_i, j, y_i, j, w_i, j, and h_i, j represent the center coordinates, width, and height of the j-th prediction box in the i-th grid cell. The <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:mover><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mover><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mrow><mml:mover><mml:mrow><mml:mtext>w</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mover><mml:mrow><mml:mtext>h</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:math></inline-formula> represent the center coordinates, width, and height of the real box, respectively. <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> indicates whether the rectangle is responsible for predicting a target object, with a value of 1 indicating responsibility and 0 indicating otherwise.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">coord</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>w</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The obj_loss was used as the confidence loss function to indicate the confidence level of the computed network. The formula for calculating obj_loss is presented in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>. The <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the BCE Loss of the confidence label matrix and the predicted confidence matrix, &#x03B1; represents the weight of confidence loss when the mask is true, usually ranging from 0.5 to 1. This makes the network focus more on training when the mask is true. Variable z represents the z-th anchor in the grid, x represents the abscissa position, y represents the ordinate position, and N represents the grid size as N&#x2009;&#x002A;&#x2009;N.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:munder><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>n</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">false</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">false</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:munder><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:msub><mml:mi>j</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The precision indicates the proportion of the predicted positive samples that are positive. The formula for calculating precision is presented in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>. The true positive (TP)&#x2009;&#x002B;&#x2009;false positive (FP) represents the number of results that have been determined to be positive samples, and TP represents the number of positive samples that have been determined to be positive.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The recall is also referred to as the check rate. It indicates the proportion of the correctly identified samples in the total positive samples. The formula for calculating recall is presented in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>. The TP&#x2009;&#x002B;&#x2009;false negative (FN) represents the actual number of positive samples, and TP represents the number of positive samples that have been determined to be positive.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The Map0.5 is the average precision of all categories at an intersection over union (IOU) threshold of 0.5. The formula for calculating Map0.5 is presented in <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. The variable k represents the total number of categories and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the average precision of the i-th category.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>Map</mml:mtext></mml:mrow><mml:mn>0.5</mml:mn><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mi>k</mml:mi></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The F1-score is a measure of the classification problem. The formula for calculating F1-score is presented in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>. The precision represents the accuracy, and recall represents the recall rate.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mtext>&#x00A0;</mml:mtext><mml:mo>&#x00D7;</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">Precision</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mo>&#x00D7;</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">Recall</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_3"><label>4.3</label><title>Performance of the TT100K Dataset</title>
<p>To test the efficacy of the proposed method, the TT100K dataset was first used to train and analyze the experimental results. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the experimental effect of the proposed method on the TT100K dataset. Given the relatively small proportion of traffic signs in the entire map, for a clearer display effect, the six localized detection detail maps were stitched together to form <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. It can be observed that the model detects targets with high accuracy and accurately classify them into the correct category.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Sample detection results based on the C2Net-YOLOv5 method</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-4.tif"/></fig>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows the results of YOLOv5 and the C2Net-YOLOv5 model trained on the TT100K dataset. The evaluation criteria consisted of cls_loss, box_loss, obj_loss, precision, recall, and Map0.5. Larger values of precision, recall, and Map0.5 indicate superior predictions, while smaller values of cls_loss, box_loss, and obj_loss denote improved predictive performance. From the comparison of the line diagram, it can be observed that in the initial stage of the detection model training, the learning efficiency of the model is higher, and the convergence rate of the training curve is faster. As the training period increases, the slope of the training curve gradually decreases and eventually stabilizes. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows that each loss function gradually converges. The precision, recall, and Map0.5 metrics continuously improve and converge as the number of training periods increases. In terms of precision metrics, the C2Net-YOLOv5 model consistently maintained a slightly higher performance than the YOLOv5 model during training. At an IOU threshold of 0.5, the proposed model achieves an improved level of detection accuracy compared to the original YOLOv5 model and converges more rapidly. Therefore, the C2Net-YOLOv5 model not only enhances precision, recall, and Map0.5, but also exhibits faster convergence than the YOLOv5 model. The C2Net-YOLOv5 model&#x2019;s performance on TT100K gradually stabilized after 600 rounds of training, indicating the feasibility of optimizing the model for significant improvements compared to the original model.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Detection results of YOLOv5 and C2Net-YOLOv5 on the TT100K dataset</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-5.tif"/></fig>
<p>In addition to the above experimental comparisons, a comparative experiment between C2Net-YOLOv5 and some other mainstream detection models was conducted. This experiment tested Faster R-CNN, SSD, YOLOv5, and C2Net-YOLOv5 on the TT100K dataset. Precision, recall, Map0.5, and F1-score were used as benchmarks for evaluation, and the detection effect of different methods are presented in <xref ref-type="table" rid="table-1">Table 1</xref>. The superiority of the proposed method compared to YOLOv5 is because of the improved ability of the backbone network to extract traffic sign features by introducing a bidirectional Res2Net module and SE attention module in YOLOv5.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Performance of the models on the TT100K dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Backbone</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">Map0.5</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Faster R-CNN</td>
<td align="left">ResNet50</td>
<td align="left">82.74</td>
<td align="left">78.67</td>
<td align="left">71.31</td>
<td align="left">80.65</td>
</tr>
<tr>
<td align="left">SSD</td>
<td align="left">VGG16</td>
<td align="left">81.35</td>
<td align="left">79.78</td>
<td align="left">72.39</td>
<td align="left">80.56</td>
</tr>
<tr>
<td align="left">YOLOv5</td>
<td align="left">CSPDarknet</td>
<td align="left">81.92</td>
<td align="left"><bold>80.31</bold></td>
<td align="left">80.05</td>
<td align="left">75.88</td>
</tr>
<tr>
<td align="left">YOLOv5_all</td>
<td align="left">C2Net</td>
<td align="left"><bold>83.74</bold></td>
<td align="left">76.31</td>
<td align="left">82.27</td>
<td align="left">79.85</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">C2Net</td>
<td align="left">83.30</td>
<td align="left">79.28</td>
<td align="left"><bold>84.23</bold></td>
<td align="left"><bold>81.24</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The reason behind the proposed method&#x2019;s superior performance over Faster R-CNN and SSD is that the feature map of Faster R-CNN is monolayered with a very small resolution. This limits its effectiveness in detecting small and multi-scale objects. In the case of SSD, its feature pyramid structure fails to harness the powerful semantic information in deep feature graphs, which is essential for effectively detecting smaller objects. Therefore, the proposed method&#x2019;s performance surpasses that of both Faster R-CNN and SSD.</p>
</sec>
<sec id="s4_4"><label>4.4</label><title>Performance of the CCTSDB 2021 Dataset</title>
<p>To demonstrate the generalization of the method, the performance of YOLOv5 and the proposed method were tested on the CCTSDB 2021 dataset. The detection results of each model on the CCTSDB 2021 dataset are presented in <xref ref-type="table" rid="table-2">Table 2</xref>. The precision, recall, Map0.5, and F1-score metrics were used as benchmarks for evaluation. The CCTSDB 2021 dataset is characterized by a small number of categories (i.e., mandatory, warning, and prohibited), and the overall performance metrics were all higher than those in the TT100K dataset. Notably, the C2Net-YOLOv5 method increased precision to 91.49, with a slight enhancement in recall, an increase in Map0.5 to 81.03, and an F1-score of 81.69. Overall, these indicators surpassed the detection effectiveness of the original YOLOv5 method.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Performance of the models on the CCTSDB 2021 dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Backbone</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">Map0.5</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">YOLOv5</td>
<td align="left">CSPDarknet</td>
<td align="left">89.06</td>
<td align="left">73.47</td>
<td align="left">79.57</td>
<td align="left">80.52</td>
</tr>
<tr>
<td align="left">YOLOv5_all</td>
<td align="left">C2Net</td>
<td align="left">86.26</td>
<td align="left">68.31</td>
<td align="left">75.21</td>
<td align="left">76.24</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">C2Net</td>
<td align="left">91.49</td>
<td align="left">73.79</td>
<td align="left">81.03</td>
<td align="left">81.69</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5"><label>4.5</label><title>Ablation Experiments</title>
<p>To gain a deeper understanding of the effect of different insertion positions of C2Net modules on the experimental results, ablation experiments were conducted. YOLOv5_all indicates that C2Net was applied to both the model&#x2019;s Backbone and Neck, while C2Net-YOLOv5 indicates that C2Net was applied solely to the Backbone segment. The key variable under consideration was whether the C2Net module for feature extraction and fusion was included. The experiments were conducted on both the TT100K and CCTSDB 2021 datasets. The impacts of the different designs on the experiments are presented in <xref ref-type="table" rid="table-3">Tables 3</xref> and <xref ref-type="table" rid="table-4">4</xref>. Upon analyzing the model&#x2019;s performance on the TT100K and CCTSDB 2021 datasets, the C2Net-YOLOv5 model exhibited the best experimental results. Notably, the recall, Map0.5, and F1-score metrics were greatly higher than the YOLOv5_all method in the TT100K dataset, despite exhibiting slightly lower accuracy. Furthermore, in the CCTSDB 2021 dataset, all metrics of C2Net-YOLOv5 surpassed the YOLOv5_all method.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Results of ablation experiments on the TT100K dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Ablation study</th>
<th align="left">Backbone</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">Map0.5</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">YOLOv5_all</td>
<td align="left">C2Net</td>
<td align="left"><bold>83.74</bold></td>
<td align="left">76.31</td>
<td align="left">82.27</td>
<td align="left">79.85</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">C2Net</td>
<td align="left">83.30</td>
<td align="left"><bold>79.28</bold></td>
<td align="left"><bold>84.23</bold></td>
<td align="left"><bold>81.24</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-4"><label>Table 4</label><caption><title>Results of ablation experiments on the CCTSDB 2021 dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Ablation study</th>
<th align="left">Backbone</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">Map0.5</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">YOLOv5_all</td>
<td align="left">C2Net</td>
<td align="left">86.26</td>
<td align="left">68.31</td>
<td align="left">75.21</td>
<td align="left">76.24</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">C2Net</td>
<td align="left"><bold>91.49</bold></td>
<td align="left"><bold>73.79</bold></td>
<td align="left"><bold>81.03</bold></td>
<td align="left"><bold>81.69</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_6"><label>4.6</label><title>Performance Comparison of Res2Net and Bidirectional-Res2Net on TT100K</title>
<p>To ascertain the true impact of the bidirectional Res2Net composition on feature extraction capability, a comparative experiment was conducted against the detection network composed of unidirectional Res2Net. The experimental comparison was conducted on the TT100K dataset, with other variables remaining constant. The C2Net-YOLOv5 model indicates that its backbone extraction module is C2Net, which uses a Bottle2neck as a bidirectional Res2Net structure. On the other hand, the C2Net-YOLOv5 model indicates that the backbone extraction module is CNet, with the only distinction being that the Bottle2neck structure changes from bidirectional to unidirectional. The impacts of the different designs on the experiments are presented in <xref ref-type="table" rid="table-5">Table 5</xref>. Upon analyzing the model&#x2019;s performance on the TT100K dataset, the C2Net-YOLOv5 model exhibited the best experimental results. All metrics of the C2Net-YOLOv5 model greatly surpassed those of the CNet-YOLOv5 method. These findings confirmed that the bidirectional Res2Net composition really enhances the feature extraction ability.</p>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Performance comparison of Res2Net and bidirectional-Res2Net on TT100K dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Backbone</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">Map0.5</th>
<th align="left">F1-score</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">CNet-YOLOv5</td>
<td align="left">CNet</td>
<td align="left">81.62</td>
<td align="left">78.89</td>
<td align="left">80.53</td>
<td align="left">79.90</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">C2Net</td>
<td align="left"><bold>83.30</bold></td>
<td align="left"><bold>79.28</bold></td>
<td align="left"><bold>84.23</bold></td>
<td align="left"><bold>81.24</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_7"><label>4.7</label><title>Overall Performance Comparison</title>
<p>This section compares the C2Net-YOLOv5 and YOLOv5_all with other methods. The results of the comparison, including input_size, Map0.5, frames per second (FPS), and speed/ms, are presented in <xref ref-type="table" rid="table-6">Table 6</xref>. It can be observed that the C2Net-YOLOv5 method exhibited even better results. Notably, Faster R-CNN achieves a Map0.5 of 82.74 for an image size of 224&#x2009;&#x00D7;&#x2009;224, but operated at a slower speed of 7 FPS. The SSD achieves a Map0.5 of 81.35 for an image size of 512&#x2009;&#x00D7;&#x2009;512, which is faster but with slightly lower detection accuracy than the proposed method. In contrast, the proposed method improves Map0.5 to 83.30 with an input image size of 640&#x2009;&#x00D7;&#x2009;640 and achieves a speed of 27.10&#x2005;FPS. Furthermore, the detection speed is 36.9&#x2005;ms per image, notably shorter than that of the two-stage detector and an improved accuracy compared to the one-stage detector SSD. This achievement highlights the proposed method&#x2019;s ability to improve the balance between accuracy and speed.</p>
<table-wrap id="table-6"><label>Table 6</label><caption><title>Comparison of different methods on the TT100K dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Method</th>
<th align="left">Input_size</th>
<th align="left">Map0.5</th>
<th align="left">FPS <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>f</mml:mtext></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:mtext>s</mml:mtext></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th align="left">Speed/ms</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Faster R-CNN</td>
<td align="left">224&#x2009;&#x00D7;&#x2009;224</td>
<td align="left">82.74</td>
<td align="left">7</td>
<td align="left">142.86&#x2005;</td>
</tr>
<tr>
<td align="left">SSD</td>
<td align="left">512&#x2009;&#x00D7;&#x2009;512</td>
<td align="left">81.35</td>
<td align="left">27.61</td>
<td align="left">36.22&#x2005;</td>
</tr>
<tr>
<td align="left">YOLOv5</td>
<td align="left">640&#x2009;&#x00D7;&#x2009;640</td>
<td align="left">81.92</td>
<td align="left">27.25</td>
<td align="left">36.7&#x2005;</td>
</tr>
<tr>
<td align="left">YOLOv5_all</td>
<td align="left">640&#x2009;&#x00D7;&#x2009;640</td>
<td align="left">83.74</td>
<td align="left">14.12</td>
<td align="left">70.8&#x2005;</td>
</tr>
<tr>
<td align="left">C2Net-YOLOv5</td>
<td align="left">640&#x2009;&#x00D7;&#x2009;640</td>
<td align="left">83.30</td>
<td align="left">27.10</td>
<td align="left">36.9&#x2005;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_8"><label>4.8</label><title>Robustness Testing</title>
<p>To evaluate the performance and reliability of the system or model in the face of various abnormal scenarios, robustness testing was conducted on the model. In practical applications, systems or models encounter various uncertainties and changes, such as changes in occlusion, target scale, and lighting. Robustness testing serves as a means to evaluate the performance of a system or model in the real world, ensuring that it adapts to various changes and diversity.</p>
<p>The robustness of the model was tested under occlusion, multi-scale changes, and noise conditions, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. The upper left image shows that the C2Net-YOLOv5 model can detect traffic signs accurately with high confidence even when more than half of the sign is obscured. This performance stems from the model training on a more diverse dataset, which includes traffic sign samples under various occlusion scenarios. This comprehensive training process enables the model to learn more robust feature representations, thereby enhancing its detection accuracy. The two comparison images in the upper right (from left to right) represent the detection performance of the original YOLOv5 and C2Net-YOLOv5 models. In the TT100K dataset, the traffic signs are often small and account for less than one percent of the entire image; the training dataset only includes small object detection. The original YOLOv5 encountered difficulties detecting large traffic signs, as it exhibited multiple detection boxes around the target and traffic signs absent in the false detection dataset. In contrast, the proposed model accurately detects large-scale traffic signs without erroneously detecting signs that are absent. Notably, the C2Net-YOLOv5 model adopts multi-scale feature extraction technology, which makes the model more effective in detecting multi-scale traffic signs. The subsequent two comparative images in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> show the model&#x2019;s accuracy in light and shadow noise scenarios. This accuracy results from the integration of HSV (hue, saturation, value) data enhancement and SE attention mechanism.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Robustness testing of the C2Net-YOLOv5 on the TT100K dataset</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42224-fig-6.tif"/></fig>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Conclusion</title>
<p>This study proposed a novel traffic sign detection method termed C2Net-YOLOv5. To address the existing limitations of YOLOv5, a bidirectional Res2Net module was used within the Bottle2neck architecture to enhance feature representation and improve the fusion of features across different scales. Additionally, the inclusion of channel attention through the SE attention mechanism within the C2Net module enabled the refining and retaining of valuable features while discarding extraneous features. The experimental results revealed that the proposed method had superior performance compared to other mainstream detection models. The proposed method effectively used feature information from different levels to improve the accuracy and robustness of object detection. Furthermore, it efficiently addressed challenges such as occlusion, changes in target scale, and lighting variations. This study has a promising impact on practical applications such as autonomous driving, video surveillance, and intelligent assistance. However, the proposed method is faced with challenges due to the common problem of deep learning: the vulnerability to malicious attacks. Future optimization works will focus on enhancing the robustness of the detection network and strengthening its ability to resist malicious adversarial attacks, thereby improving application security. Some possible optimization directions include adversarial attack training, robustness enhancement technology, detection of abnormal input, data enhancement, and preprocessing.</p>
</sec>
</body>
<back>
<ack>
<p>We would like to express our sincere appreciation to the National Key R&#x0026;D Program of China and the Beijing Natural Science Foundation for providing the necessary financial support to conduct this research project.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research was funded by the National Key R&#x0026;D Program of China, Grant Number 2017YFB0802803, Beijing Natural Science Foundation, Grant Number 4202002.</p></sec>
<sec><title>Author Contributions</title>
<p>Study conception and design: Xiujuan Wang, Yiqi Tian; data collection: Kangfeng Zheng, Chutong Liu; analysis and interpretation of results: Xiujuan Wang, Yiqi Tian; draft manuscript preparation: Yiqi Tian. All authors reviewed the results and approved the final version of the manuscript.</p></sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available.</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. H.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Cheng</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>X. Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Yang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Res2Net: A new multi-scale backbone architecture</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>43</volume>, no. <issue>2</issue>, pp. <fpage>652</fpage>&#x2013;<lpage>662</lpage>, <year>2021</year>; <pub-id pub-id-type="pmid">31484108</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>C. S.</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A multiscale recognition method for the optimization of traffic signs using GMM and category quality focal loss</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>20</volume>, no. <issue>17</issue>, pp. <fpage>4850</fpage>&#x2013;<lpage>4850</lpage>, <year>2020</year>; <pub-id pub-id-type="pmid">32867246</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>A novel traffic sign detection method via color segmentation and robust shape matching</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>169</volume>, no. <issue>2</issue>, pp. <fpage>77</fpage>&#x2013;<lpage>88</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Maldonado-Basc&#x00F3;n</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Lafuente-Arroyo</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Gil-Jim&#x00E9;nez</surname></string-name>, <string-name><given-names>H.</given-names> <surname>G&#x00F3;mez-Moreno</surname></string-name> and <string-name><given-names>F.</given-names> <surname>L&#x00F3;pez-Ferreras</surname></string-name></person-group>, &#x201C;<article-title>Road-sign detection and recognition based on support vector machines</article-title>,&#x201D; <source>IEEE Transactions on Intelligent Transportation Systems</source>, vol. <volume>8</volume>, no. <issue>2</issue>, pp. <fpage>264</fpage>&#x2013;<lpage>278</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Donahue</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Darrell</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Malik</surname></string-name></person-group>, &#x201C;<article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Columbus, OH, USA</conf-loc>, pp. <fpage>580</fpage>&#x2013;<lpage>587</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name></person-group>, &#x201C;<article-title>Fast R-CNN</article-title>,&#x201D; in <conf-name>Proc. of IEEE Int. Conf. on Computer Vision (ICCV)</conf-name>, <conf-loc>Santiago, Chile</conf-loc>, pp. <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Faster R-CNN: Towards real-time object detection with region proposal networks</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>39</volume>, no. <issue>6</issue>, pp. <fpage>1137</fpage>&#x2013;<lpage>1149</lpage>, <year>2017</year>; <pub-id pub-id-type="pmid">27295650</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T. Y.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Doll&#x00E1;r</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name>, <string-name><given-names>K. M.</given-names> <surname>He</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Hariharan</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Feature pyramid networks for object detection</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, pp. <fpage>936</fpage>&#x2013;<lpage>944</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>H. F.</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>J. P.</given-names> <surname>Shi</surname></string-name> and <string-name><given-names>J. Y.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>Path aggregation network for instance segmentation</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Salt Lake City, USA</conf-loc>, pp. <fpage>8759</fpage>&#x2013;<lpage>8768</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z. W.</given-names> <surname>Cai</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Vasconcelos</surname></string-name></person-group>, &#x201C;<article-title>Cascade R-CNN: Delving into high quality object detection</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Salt Lake City, USA</conf-loc>, pp. <fpage>6154</fpage>&#x2013;<lpage>6162</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>G. Y.</given-names> <surname>Gao</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Real-time small traffic sign detection with revised Faster-RCNN</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>78</volume>, no. <issue>10</issue>, pp. <fpage>13263</fpage>&#x2013;<lpage>13278</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Redmon</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Divvala</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Farhadi</surname></string-name></person-group>, &#x201C;<article-title>You only look once: Unified, real-time object detection</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Las Vegas, NV, USA</conf-loc>, pp. <fpage>779</fpage>&#x2013;<lpage>788</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Anguelov</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Erhan</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Szegedy</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Reed</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>SSD: Single shot multiBox detector</article-title>,&#x201D; in <conf-name>Proc. of Computer Vision-ECCV 2016</conf-name>, <conf-loc>Amsterdam, Netherlands</conf-loc>, pp. <fpage>21</fpage>&#x2013;<lpage>37</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Redmon</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Farhadi</surname></string-name></person-group>, &#x201C;<article-title>YOLOv3: An incremental improvement</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Salt Lake City, USA</conf-loc>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>You</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Ke</surname></string-name>, <string-name><given-names>H. P.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W. H.</given-names> <surname>You</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Wu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Small traffic sign detection and recognition in high-resolution images</article-title>,&#x201D; in <conf-name>Proc. of Int. Conf. on Cognitive Computing</conf-name>, <conf-loc>Sichuan, China</conf-loc>, pp. <fpage>37</fpage>&#x2013;<lpage>53</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kong</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Jang</surname></string-name></person-group>, &#x201C;<article-title>Lightweight traffic sign recognition algorithm based on cascaded CNN</article-title>,&#x201D; in <conf-name>Proc. of Int. Conf. on Control, Automation and Systems (ICCAS)</conf-name>, <conf-loc>Jeju, Korea (South)</conf-loc>, pp. <fpage>506</fpage>&#x2013;<lpage>509</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. X.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>X. H.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Brandt</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Hua</surname></string-name></person-group>, &#x201C;<article-title>A convolutional neural network cascade for face detection</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Boston, MA, USA</conf-loc>, pp. <fpage>1063</fpage>&#x2013;<lpage>6919</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Redmon</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Farhadi</surname></string-name></person-group>, &#x201C;<article-title>YOLO9000: Better, faster, stronger</article-title>,&#x201D; in <conf-name>Proc. of IEEE Conf. on Computer Vision and Pattern Recognitio (CVPR)</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, pp. <fpage>6517</fpage>&#x2013;<lpage>6525</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. H.</given-names> <surname>Yen</surname></string-name>, <string-name><given-names>C. Y.</given-names> <surname>Shu</surname></string-name> and <string-name><given-names>H. H.</given-names> <surname>Hsu</surname></string-name></person-group>, &#x201C;<article-title>Occluded traffic signs recognition</article-title>,&#x201D; in <conf-name>Proc. of Future of Information and Communication Conf.</conf-name>, <conf-loc>San Francisco, California, USA</conf-loc>, pp. <fpage>794</fpage>&#x2013;<lpage>804</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Siniosoglou</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Sarigiannidis</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Spyridis</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Khadka</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Efstathopoulos</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Synthetic traffic signs dataset for trafficsign detection &#x0026; recognition in distributedsmart systems</article-title>,&#x201D; in <conf-name>Proc. of Int. Conf. on Distributed Computing in Sensor Systems (DCOSS)</conf-name>, <conf-loc>Pafos, Cyprus</conf-loc>, pp. <fpage>102692</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Franzen</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Yuan</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Impact of frequency-domain data representation on recognition performance of neural networks</article-title>,&#x201D; <source>Developments of Artificial Intelligence Technologies in Computation and Robotics</source>, vol. <volume>12</volume>, no. <issue>12015</issue>, pp. <fpage>1188</fpage>&#x2013;<lpage>1195</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Jocher</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Stoken</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Borovec</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Chaurasia</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Liu</surname></string-name> <etal>et al.,</etal></person-group> <source>Ultralytics/yolov5: v5.0-YOLOv5-P6 1280 models, AWS, Supervise.ly and YouTube integrations</source>, <comment>Zenodo</comment>, <year>2021</year>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://api.semanticscholar.org/CorpusID:244964519">https://api.semanticscholar.org/CorpusID:244964519</ext-link></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Nagrath</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Jain</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Madan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Arora</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Kataria</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>SSDMNV2: A real time DNN-based face mask detection system using single shot multibox detector and MobileNetV2</article-title>,&#x201D; <source>Sustainable Cities and Society</source>, vol. <volume>66</volume>, no. <issue>5</issue>, pp. <fpage>102692</fpage>, <year>2021</year>; <pub-id pub-id-type="pmid">33425664</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Pooja</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Preeti</surname></string-name></person-group>, &#x201C;<chapter-title>Face mask detection using AI</chapter-title>,&#x201D; in <source>Predictive and Preventive Measures for COVID-19 Pandemic</source>, vol. <volume>1</volume>, <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>293</fpage>&#x2013;<lpage>305</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yun</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Manipulator grabbing position detection with information fusion of color image and depth image using deep learning</article-title>,&#x201D; <source>Journal of Ambient Intelligence and Humanized Computing</source>, vol. <volume>12</volume>, no. <issue>12</issue>, pp. <fpage>10809</fpage>&#x2013;<lpage>10822</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>M-YOLO: Traffic sign detection algorithm applicable to complex scenarios</article-title>,&#x201D; <source>Symmetry</source>, vol. <volume>14</volume>, no. <issue>5</issue>, pp. <fpage>952</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Loey</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Manogaran</surname></string-name>, <string-name><given-names>M. H. N.</given-names> <surname>Taha</surname></string-name> and <string-name><given-names>N. E. M.</given-names> <surname>Khalifa</surname></string-name></person-group>, &#x201C;<article-title>A hybrid deep transfer learning model with machine learning methods for face mask detection in the era of the COVID-19 pandemic</article-title>,&#x201D; <source>Measurement</source>, vol. <volume>167</volume>, no. <issue>1</issue>, pp. <fpage>108288</fpage>, <year>2021</year>; <pub-id pub-id-type="pmid">32834324</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yun</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Tao</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Grasping pose detection for loose stacked object based on convolutional neural network with multiple self-powered sensors information</article-title>,&#x201D; <source>IEEE Sensors Journal</source>, vol. <volume>23</volume>, no. <issue>18</issue>, pp. <fpage>20619</fpage>&#x2013;<lpage>20632</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. L.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>X. D.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>M. M.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>SC-YOLO: A object detection model for small traffic signs</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>11</volume>, no. <issue>1</issue>, pp. <fpage>11500</fpage>&#x2013;<lpage>11510</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Jiang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Deep learning based 3D target detection for indoor scenes</article-title>,&#x201D; <source>Applied Intelligence</source>, vol. <volume>53</volume>, no. <issue>9</issue>, pp. <fpage>10218</fpage>&#x2013;<lpage>10231</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. M.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>L. D.</given-names> <surname>Kuang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>R. S.</given-names> <surname>Sherratt</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>CCTSDB 2021: A more comprehensive traffic sign detection benchmark</article-title>,&#x201D; <source>Human-Centric Computing and Information Sciences</source>, vol. <volume>12</volume>, no. <issue>23</issue>, pp. <fpage>13263</fpage>&#x2013;<lpage>13278</lpage>, <year>2021</year>.</mixed-citation></ref>
</ref-list>
</back></article>