<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">39451</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.039451</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Lightweight Surface Litter Detection Algorithm Based on Improved YOLOv5s</article-title>
<alt-title alt-title-type="left-running-head">Lightweight Surface Litter Detection Algorithm Based on Improved YOLOv5s</alt-title>
<alt-title alt-title-type="right-running-head">Lightweight Surface Litter Detection Algorithm Based on Improved YOLOv5s</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Zunliang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Chengxu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Duan</surname><given-names>Lucheng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Tan</surname><given-names>Baohua</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>tbh@hbut.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Science (College of Chip Industry), Hubei University of Technology</institution>, <addr-line>Wuhan, 430068</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>National &#x201C;111 Research Center&#x201D; Microelectronics and Integrated Circuits, Hubei University of Technology</institution>, <addr-line>Wuhan, 430068</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Baohua Tan. Email: <email>tbh@hbut.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2023</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>09</day>
<month>6</month>
<year>2023</year></pub-date>
<volume>76</volume>
<issue>1</issue>
<fpage>1085</fpage>
<lpage>1102</lpage>
<history>
<date date-type="received"><day>30</day><month>1</month><year>2023</year></date>
<date date-type="accepted"><day>17</day><month>4</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Chen et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Chen et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_39451.pdf"></self-uri>
<abstract>
<p>In response to the problem of the high cost and low efficiency of traditional water surface litter cleanup through manpower, a lightweight water surface litter detection algorithm based on improved YOLOv5s is proposed to provide core technical support for real-time water surface litter detection by water surface litter cleanup vessels. The method reduces network parameters by introducing the deep separable convolution GhostConv in the lightweight network GhostNet to substitute the ordinary convolution in the original YOLOv5s feature extraction and fusion network; introducing the C3Ghost module to substitute the C3 module in the original backbone and neck networks to further reduce computational effort. Using a Convolutional Block Attention Mechanism (CBAM) module in the backbone network to strengthen the network&#x2019;s ability to extract significant target features from images. Finally, the loss function is optimized using the Focal-EIoU loss function to improve the convergence speed and model accuracy. The experimental results illustrate that the improved algorithm outperforms the original Yolov5s in all aspects of the homemade water surface litter dataset and has certain advantages over some current mainstream algorithms in terms of model size, detection accuracy, and speed, which can deal with the problems of real-time detection of water surface litter in real life.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Surface litter detection</kwd>
<kwd>lightweight</kwd>
<kwd>YOLOv5s</kwd>
<kwd>GhostNet</kwd>
<kwd>deep separable convolution</kwd>
<kwd>convolutional block attention mechanism (CBAM)</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>China University Industry-University Research Innovation Fund Project</funding-source>
<award-id>2022BL052</award-id>
</award-group>
<award-group id="awg2">
<funding-source>the Science and Technology Innovation R&#x0026;D Project of the State General Administration of Sports of China</funding-source>
<award-id>22KJCX024</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Major Project of Philosophy and Social Science Research in Higher Education Institutions in Hubei Province</funding-source>
<award-id>21ZD054</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Key Project of Hubei Provincial Key Laboratory of Intelligent Transportation Technology and Equipment Open Fund</funding-source>
<award-id>2022XZ106</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Along with the high-quality growth of China&#x2019;s economy and the rising living standards of its residents, the increasing richness of material life is accompanied by the corresponding phenomena of the massive output of garbage, random discarding, simple piling, and disposal [<xref ref-type="bibr" rid="ref-1">1</xref>]. Besides, the problems of water pollution and eutrophication caused by floating litter on the water surface have seriously affected the ecological civilization and human living environment in the watershed. At the present stage, the management of floating litter on water surfaces by domestic related departments and institutions is mainly based on interception and collection and manual salvage, but the manual cleaning method is time-consuming, laborious, and long-period, and it cannot reach the demands of intelligent detection and automatic cleaning in real-time.</p>
<p>With modern information technology and the advancement of intelligent manufacturing technology, related researchers have proposed the combination of object detection technology and mechanical cleanup boats to detect and clean up floating litter on the water surface by using deep learning technology instead of manual labor [<xref ref-type="bibr" rid="ref-2">2</xref>]. Since these scenes use hardware that is mostly edge devices, under the terms of limited memory and arithmetic power, it is necessary to take into account not only its mechanical cleanup device drive system requirements, the required arithmetic power requirements of the model. Therefore, reducing the model size is more beneficial for porting to computing devices [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>As artificial intelligence technology and computer vision develop, object detection, as one of the key branches, is extensively applied in Optical Character Recognition (OCR) analysis, contextual analysis, disaster management, and vehicle recognition. In the field of Unmanned Aerial Vehicles (UAV), Dilshad et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] innovated the LocateUAV method for estimating UAV location using contextual analysis in an Internet of Things (IoT) environment. In the vision field, the traditional object detection technology based on machine learning mainly uses the size, shape, color, texture, and other information of the object, and then performs template matching and saliency detection after completing manual feature extraction by image segmentation technology [<xref ref-type="bibr" rid="ref-5">5</xref>]. However, this method suffers from insufficient model detection accuracy, speed, and poor robustness. Over the past few years, object detection techniques grounded on deep learning have obtained outstanding results in terms of detection effectiveness and model robustness under their convolutional neural networks&#x2019; ability to automatically extract features. Yu et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] proposed the High-resolution Network (HRNet) improves detection accuracy by fusing multi-scale feature maps using parallel structural networks to obtain different multi-scale feature information. Tang et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] achieved good detection results in OCR by changing the detection scale of the network and incorporating a modified OCR branch in the network. Currently, there are two stages in the development of object recognition algorithms based on deep learning: one class is region-based two-stage object detection algorithms, with networks such as Regions Convolutional Neural Network (R-CNN) [<xref ref-type="bibr" rid="ref-8">8</xref>], Fast R-CNN [<xref ref-type="bibr" rid="ref-9">9</xref>], and Faster R-CNN [<xref ref-type="bibr" rid="ref-10">10</xref>] as typical representatives. This type of method performs feature and classification detection in two stages based on the divided regions, which have a higher detection accuracy but slow speed. Another class is the single-stage regression-based object detection algorithm, with a series of networks such as You Only Look Once (YOLO) [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>] and Single Shot Multi-Box Detector (SSD) [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>] as typical representatives. Mathias et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] merged a Gaussian mixture model 2D empirical modal decomposition algorithm with a Yolov3 depth network and applied it to underwater object detection.</p>
<p>The YOLO series algorithm integrates features and classification into one network and identifies and locates them through regression calculation, significantly increasing the detection speed and meeting the demand for detection in real-time, but with a loss in detection accuracy. For that reason, this paper proposed a lightweight surface litter detection algorithm with improved YOLOv5s, which reduces the model size and makes the network more lightweight based on improving the detection performance of the original YOLOv5s methods. The main contributions of our work are as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>Improving the ordinary convolution in the YOLOv5s network into a deeply separable convolution GhostConv, which effectively reduces the convolution layers and computational resources; introducing lightweight C3Ghost modules to substitute the C3 module, further minimizing the model parameters and computation, making it more lightweight.</p></list-item>
<list-item><label>(2)</label><p>This paper fuses the CBAM module [<xref ref-type="bibr" rid="ref-16">16</xref>] in the backbone network of YOLOv5s, applying it to enrich the network&#x2019;s capacity to acquire target feature information and improve the model&#x2019;s prediction capability.</p></list-item>
<list-item><label>(3)</label><p>A new loss function is adopted, and the Focal-EIoU (Focal and Efficient Intersection over Union) [<xref ref-type="bibr" rid="ref-17">17</xref>] loss function is invoked as the prediction frame regression loss function in the YOLOv5s network to optimize the deficiencies of the CIoU (Distance IoU) [<xref ref-type="bibr" rid="ref-18">18</xref>] loss function in the original method and accelerate the loss function convergence speed and model prediction accuracy.</p></list-item>
</list></p>
<p>In this paper, Section 2 reviews the work on water surface litter detection, YOLOv5s network, and lightweight network. Section 3 describes the water surface litter detection methods. Section 4 describes the experimental content, including the experimental datasets production, experimental environment, and evaluation metrics. Experiments and results are presented and reviewed in Section 5. Section 6 summarizes and gives an outlook.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<sec id="s2_1"><label>2.1</label><title>Water Surface Detection</title>
<p>Water surface litter detection is a branch of object detection, which has significant research implications for environmental management and water resource protection. Surface object detection has been gradually improved and enriched from traditional manually designed features and shallow classifier frameworks based on deep learning object detection frameworks.</p>
<p>As far as the traditional detection methods are concerned, they rely too much on key points, edges, and templates leading to their low detection accuracy. For example, Matsumoto&#x2019;s Histogram of Oriented Gradient (HOG)-Support Vector Machine (SVM) method [<xref ref-type="bibr" rid="ref-19">19</xref>], proposed in 2013, detects surface vessels by images taken by a shipboard camera. Kaido et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] combined SVM methods and edge detection techniques for ship detection and ship number identification. However, a majority of these methods traditionally scan the entire image through a swiping window to detect objects, which greatly limits detection efficiency.</p>
<p>Deep learning techniques, with their powerful image processing power and expression performance of convolutional neural networks, are applied to obtain remarkable effects in surface object detection. Zhang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] optimized the Faster R-CNN by fusing high-level and low-level features of the network to enrich the real-time detection of surface objects. Panwar et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed the AquaTrash dataset using deep migration learning techniques to enrich the generalizability of the AquaTrash dataset. Li et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] improved the YOLOv3 algorithm and embedded it into a water surface garbage cleaning robot only to achieve intelligent cleaning of water surface litter. The rich surface object dataset and constantly innovative and iterative object detection algorithms provide key technical support for surface object detection.</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>YOLOv5 Method</title>
<p>The YOLOv5 [<xref ref-type="bibr" rid="ref-24">24</xref>] network model is made up of four main parts: Input, Backbone, Neck, and Head network, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Input side uses Mosaic&#x2009;&#x002B;&#x2009;Mixup data enhancement technology to randomly crop, scale, and stitch photos from the input into enroll the background of the dataset images, enhance the network&#x2019;s generalizability. The Backbone network mainly consists of Cross Stage Partial (CBS), CSP Bottleneck with 3 convolutions (C3), and Spatial Pyramid Pooling-Fast (SPPF) modules, where the CBS structure consists of the Conv&#x2009;&#x002B;&#x2009;BN&#x2009;&#x002B;&#x2009;SiLU activation function, which convolves and normalizes the input images before passing them through the activation function to the next layer of convolution. The C3 structure is a reworked version of the CSP structure made in the correction unit, for dividing the feature mappings at the base layer before merging them via the inter-stage hierarchy to ensure the correct rate while reducing the computational bottleneck and strengthening the network&#x2019;s learning capability. The SPPF module uses three 5&#x2009;&#x00D7;&#x2009;5 maximum pooling to fuse feature map messages across different scales to enrich the network&#x2019;s ability to abstract features from images. The neck network uses Feature Pyramid Network (FPN)&#x2009;&#x002B;&#x2009;Path Aggregation Network (PAN) [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>] structure to further enhance the network feature fusion capability using a combination of top-down and bottom-up approaches. The head network screens the target candidate frames by non-maximal suppression and is used as the output of the prediction results.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>YOLOv5s overall network design chart</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-1.tif"/></fig>
</sec>
<sec id="s2_3"><label>2.3</label><title>Lightweight Network</title>
<p>A lightweight network means that the size and parameters of the model are reduced as much as possible by optimizing the network structure to meet the requirements of low arithmetic power of edge devices while ensuring that the model is equally effective in detection. At present, most of the more popular lightweight solutions of the algorithm are considered from model compression and pattern structure design.</p>
<p>Model reduction aims to take an already trained network and reduce the number of parameters, usually with schemes such as model distillation, pruning, and quantization. With the improvement in the level of hardware conditions, the computational power of computers has grown by leaps and bounds, and Neural Architecture Search (NAS) network is also a technique to find the best network using computational-level arithmetic power [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<p>Among the model architecture design solutions, many excellent design solutions have been broadly applied in the past few years. The more classical lightweight network structures are listed from the time of computational introduction, such as SqueezeNet [<xref ref-type="bibr" rid="ref-29">29</xref>], which proposes Fire Module, ShuffleNet [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>] series of ShuffleNet Unit feature fusion schemes, etc.</p>
<p>Google proposed MobileNetV1 [<xref ref-type="bibr" rid="ref-32">32</xref>] object detection algorithm using a depth-separable convolution structure and in the subsequent proposed MobileNetV2 [<xref ref-type="bibr" rid="ref-33">33</xref>] and MobileNetV3 [<xref ref-type="bibr" rid="ref-34">34</xref>] algorithms, the essence of depth-separable convolution involves splitting the normal convolution into deep and pointwise convolution [<xref ref-type="bibr" rid="ref-35">35</xref>]. In particular, in terms of network structure, MobileNetV3 derives the main network structure through a neural architecture search [<xref ref-type="bibr" rid="ref-36">36</xref>]. The network uses a 5&#x2009;&#x00D7;&#x2009;5 depth-separable convolutional structure by adding the Squeeze-and-Excitation (SE) attention mechanism [<xref ref-type="bibr" rid="ref-37">37</xref>], which assigns its weights on the feature maps through the process of network training, which is very easy to use, plug-and-play, and has shown good performance improvement in several mainstream networks.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Water Surface Litter Detection Method</title>
<p>For water surface litter detection, the main detection process can be divided into three steps: dataset production, model improvement, and model deployment. Since there is no publicly available water surface litter dataset, the first step requires a homemade water surface litter dataset. After filtering the collected images, the object garbage in the pictures is labeled with Labelimg labeling software, with the labeled garbage categories divided into 7 categories, namely: &#x007B;lunch boxes, foam, bottles, plastic bags, leaves and branches, food packaging bags, paper drink boxes&#x007D;, which are used to generate a standard VOC dataset format. Labelimg labeling software will automatically label the object boxes with the labeled information the generated XML file is stored in the Annotations folder, the TXT file of the dataset division is stored in the ImageSets folder, and the JPEGImages are used to store the original images. In the second step, the produced dataset is put into the enhanced YOLOv5s network for training, and the network will perform Mosaic&#x2009;&#x002B;&#x2009;Mixup [<xref ref-type="bibr" rid="ref-38">38</xref>] data enhancement process on the dataset according to the pre-set ratio, and the CBAM module is inserted in the bottom position of the improved YOLOv5s backbone network to help extract features. It is then fed into the feature fusion network for training after the improved Focal-EIoU loss function balances the samples of positive and negative. Finally, to evaluate the trained model with test data, and after completing the evaluation, it is deployed into the embedded device on the water surface trash cleaning vessel to provide real-time detection and intelligent classification of water surface trash, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Detection flow chart</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-2.tif"/></fig>
<sec id="s3_1"><label>3.1</label><title>Improving Convolution and C3 Module of YOLOv5 Network</title>
<p>YOLOv5 [<xref ref-type="bibr" rid="ref-24">24</xref>] network is an open-source object detection algorithm proposed by Ultralytics and is the fifth edition of the YOLO series developed to date. A typical single-stage object detection algorithm, YOLOv5 is divided into four versions in order of detection accuracy and model size, depending on network layers&#x2019; depth and feature map width: YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x, which they respond to the different demands of industrial applications concerning detection accuracy, detection speed, and the module size, respectively. In this article, YOLOv5s is chosen as the benchmark model from the perspective of edge devices oriented to low computing power.</p>
<p>The network structure diagram of YOLOv5s shows that the network contains many basic convolutional blocks of CBS are made up of Convolution (Conv), Batch Normalization (BN), Sigmoid Linear Unit (SiLU), and C3 structure consisting of CBS, Bottleneck, and Concat. More basic convolutional blocks increase network computational parameters and affect network inference speed. Therefore, in this paper, from the perspective of lightweight, a more concise network structure is used to replace the more computationally intensive ordinary convolutional Conv and C3 structures in YOLOv5s based on the premise of equal detection effect. The ordinary convolutional Conv in the YOLOv5s network is substituted by the deep separable convolutional GhostConv, as displayed in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>CBS comparison chart</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-3.tif"/></fig>
<p>GhostNet [<xref ref-type="bibr" rid="ref-39">39</xref>] network is an efficient network designed by Huawei Noah&#x2019;s Ark Lab, whose core idea is to use less computationally intensive Ghost modules to generate redundant feature maps that exist in neural networks. A feature of the Ghost module is to instead of the normal convolutional layer, it divides the normal convolution into two parts, which are divided into two parts. First, the necessary feature condensation map of the input features is obtained using the ordinary 1&#x2009;&#x00D7;&#x2009;1 convolution that acts similarly to feature integration; then the feature con map obtained in the previous stage to obtain the similar character map of feature condensation using the depth-separable convolution. By this two-step operation, the number of model counts is reduced by obtaining redundant characteristic maps while minimizing the number of convolutional layers, as displayed in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Schematic diagram of the two convolution processes, (a) normal convolution (b) deeply separable convolution GhostConv</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-4.tif"/></fig>
<p>The GhostNet network structure is composed of a GhostNet Bottlenecks section as the backbone and a residual edge section. Bottlenecks mainly consist of a bottleneck structure with two stacked Ghost modules. Add channels and set the expansion ratio of the number of output channels with the first ghost module; the second ghost module decreases the channel count to match the number of input channels. This structure can be categorized into two types, depending on the step size. A deep separable convolution with Stride&#x2009;&#x003D;&#x2009;2 is used for twice the down-sampling. In this paper, the C3 module in the YOLOv5s network is substituted with the C3Ghost module, which decreases and makes the overall YOLOv5s model more lightweight and upgrades the execution speed of the model. The new network is displayed in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Structure of the improved C3Ghost network</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-5.tif"/></fig>
</sec>
<sec id="s3_2"><label>3.2</label><title>Introduces CBAM Module</title>
<p>The CBAM module is a lightweight attention network presented by Woo et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] that enhances the feature representation of the network with attention operations that can be performed in the channel and spatial dimension. In this article, the CBAM module is integrated before the SPPF module in the YOLOv5s backbone network to strengthen the extraction capability of small target feature information using the CBAM module. The modified YOLOv5s network structure is displayed in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> below.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Improved YOLOv5s network structure</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-6.tif"/></fig>
<p>The CBAM module contains two independent sub-modules, Channel Attention Module (CAM) and Spatial Attention Module (SAM), and the overall structure is presented in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. A feature map <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>F</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of an intermediate layer is given as input, a 1D channel attention feature map <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and a 2D spatial attention feature map <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, with <italic>C</italic> denoting the channel of the feature map, and <italic>H</italic> and <italic>W</italic> denoting the height and width of the feature map. The input feature map <italic>F</italic> is first multiplied with the feature map <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of <italic>F</italic> after channel attention module operation to get <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and then <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is multiplied with the feature map <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> after spatial attention module operation to get <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, and its whole calculation process is given in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2297;</mml:mo><mml:mi>F</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula> denotes multiplication by elements.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>The CBAM module structure chart</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-7.tif"/></fig>
<p>The CAM is presented in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The input <italic>F</italic> is compressed in spatial dimension by two operations of maximum pooling and average pooling to get two 1&#x2009;&#x00D7;&#x2009;1&#x2009;&#x00D7;&#x2009;<italic>C</italic> feature maps <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, which are fed into a shared network composed of Multilayer Perceptron (MLP) [<xref ref-type="bibr" rid="ref-40">40</xref>] for the calculation to obtain two different background description maps, then make it summed at the pixel level and then use a sigmoid function to activate it to finally get the channel attention map <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and the operation process is given in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">MaxPool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>C</italic> is the output vector length; <italic>&#x03C3;</italic> is the sigmoid activation function; <italic>MLP</italic> is the shared fully connected layer; <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mfrac><mml:mi>C</mml:mi><mml:mi>r</mml:mi></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denotes the first layer of the shared full connectivity layer; <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>C</mml:mi><mml:mi>r</mml:mi></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula> denotes the second layer of the shared full connectivity layer.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>Channel attention module (CAM)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-8.tif"/></fig>
<p>The SAM is given in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. Firstly, maximum pooling and average pooling are applied in the channel dimension to acquire two <italic>H&#x2009;&#x00D7;&#x2009;W&#x2009;&#x00D7;&#x2009;1</italic> feature maps, then channel cascading is performed on the feature maps, followed by a convolutional layer to downscale to a single channel, and finally the spatial attention map <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is obtained after activation by a sigmoid function, and the operation process is illustrated in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>7</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>A</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>;</mml:mo><mml:mrow><mml:mtext mathvariant="italic">MaxPol</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;&#x00A0;</mml:mtext><mml:mspace width="mediummathspace" /><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>7</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>;</mml:mo><mml:msubsup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msup><mml:mi>f</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>7</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> denotes a 7&#x2009;&#x00D7;&#x2009;7 convolutional layer.</p>
<fig id="fig-9"><label>Figure 9</label><caption><title>Spatial attention module (SAM)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-9.tif"/></fig>
</sec>
<sec id="s3_3"><label>3.3</label><title>Optimization of the Loss Function</title>
<p>The loss function of the YOLOv5s network is made up of three loss functions, which are localization loss (loss<sub>box</sub>), confidence loss (loss<sub>obj</sub>), and classification loss (loss<sub>cls</sub>). The magnitude of the loss function value is the sum of the three loss functions, as given in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>. The CIoU [<xref ref-type="bibr" rid="ref-18">18</xref>] loss function is used as the bounding box regression loss function by default in the YOLOv5s network. The formula is shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi>&#x03BD;</mml:mi></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mi>&#x03BD;</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03BD;</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>&#x03BD;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>4</mml:mn><mml:msup><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>arctan</mml:mi></mml:mrow><mml:mfrac><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mfrac><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi>arctan</mml:mi></mml:mrow><mml:mfrac><mml:mi>w</mml:mi><mml:mi>h</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>In the formula, <italic>IoU</italic> indicates how the predicted frame intersects the true frame; <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> for the calculation of a Euclidean distance between two centroids; <italic>c</italic> shows the diagonal distance between the minimum external rectangle of the object box to be detected and the real object box; <italic>&#x03B1;</italic> indicates a weighting factor; <italic>&#x03BD;</italic> is a measure of aspect ratio consistency; <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mfrac></mml:mstyle></mml:math></inline-formula> indicates the aspect ratio of the real frame; <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>w</mml:mi><mml:mi>h</mml:mi></mml:mfrac></mml:mstyle></mml:math></inline-formula> indicates the length-to-width ratio of the prediction frame.</p>
<p>While CIoU considers accounting for the intersection area, centroid distance, and aspect ratio of the prediction frame regression, the difference in aspect ratio is reflected by <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>&#x03BD;</mml:mi></mml:math></inline-formula> in the formula and is not the real difference between the width and height respectively, and its confidence level, which sometimes prevents effective optimization of the model. EIoU [<xref ref-type="bibr" rid="ref-17">17</xref>] splits the aspect ratio based on CIoU to explicitly measure the difference between the three geometric factors, accelerating convergence and improving regression accuracy. Simultaneously, Focal loss is added to tweak the problem of unbalanced hard and easy samples. Therefore, in this paper, Focal-EIoU [<xref ref-type="bibr" rid="ref-17">17</xref>] is introduced as the prediction frame regression loss function in the YOLO v5s network, and a suppression factor <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is added, which is calculated as shown in <xref ref-type="disp-formula" rid="eqn-8">Eqs. (8)</xref>, <xref ref-type="disp-formula" rid="eqn-9">(9)</xref>.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;&#x00A0;</mml:mtext><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext>g</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Focal</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>E</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Experimental Design</title>
<sec id="s4_1"><label>4.1</label><title>Dataset Production</title>
<p>Currently, there is no public open-source dataset for research on water surface litter classification. To evaluate the capabilities of the model, the types of surface litter were classified into seven categories: bottles, paper drink containers, lunch boxes, foam, plastic bags, food packaging bags, and leaves and branches, according to domestic litter classification standards and by combining the types, forms, and sizes of floating litter commonly found on the water surface of rivers and lakes. All data sets used in the experiments were acquired in the watershed of the XunSi River at the Hubei University of Technology. To ensure the comprehensiveness and complexity of the datasets and to enhance the generalization of the detector, the collection of photos needed to take into account different periods and weather. A total of 3000 photos in JPG format with 480&#x2009;&#x00D7;&#x2009;480 pixels were acquired. To overcome problems of overfitting caused by small data sets and the effects of water reflection and lighting on detection, image processing methods such as cropping, rotation, flipping, and stitching are used to improve the diversity of data sets and the richness of image backgrounds, which are helpful for feature extraction of small target objects. As shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<fig id="fig-10"><label>Figure 10</label><caption><title>Schematic diagram of different scenario datasets. (a) photos of water surface debris at different times (b) spliced photos</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-10.tif"/></fig>
<p>The images in the dataset are labeled according to the determined categories using the open-source software Labelimg, which generates standard VOC dataset format files, including three folders of Annotation, ImageSets, and JPEGImages, containing all the labeling information of the images. To avoid the problem of model overfitting, the model should have the best generalization performance. This paper divides the produced dataset into training and test sets at the rate of 9:1, where 2700 images are used as the training set and 300 pictures are used as the test set, and using migration training method for 100 rounds of epoch training. As presented in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Table of the count of various types of water surface litter</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Category</th>
<th align="left">Lunch boxes</th>
<th align="left">Foam</th>
<th align="left">Bottles</th>
<th align="center">Plastic bags</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Category number</td>
<td align="left">0</td>
<td align="left">1</td>
<td align="left">2</td>
<td align="left">3</td>
</tr>
<tr>
<td align="left">Training set</td>
<td align="left">1520</td>
<td align="left">1183</td>
<td>1768</td>
<td align="left">709</td>
</tr>
<tr>
<td align="left">Test set</td>
<td align="left">166</td>
<td align="left">125</td>
<td align="left">218</td>
<td align="left">76</td>
</tr>
<tr>
<td align="left">Category</td>
<td align="left">Leaves and branches</td>
<td align="left">Food packaging bags</td>
<td align="left" colspan="2">Paper drink boxes</td>
</tr>
<tr>
<td align="left">Category number</td>
<td align="left">4</td>
<td align="left">5</td>
<td align="left" colspan="2">6</td>
</tr>
<tr>
<td align="left">Training set</td>
<td align="left">1868</td>
<td align="left">1060</td>
<td align="left" colspan="2">1394</td>
</tr>
<tr>
<td align="left">Test set</td>
<td align="left">212</td>
<td align="left">86</td>
<td align="left" colspan="2">163</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2"><label>4.2</label><title>Experimental Configuration</title>
<p>In this paper, the experimental environment for the Windows 10 operating system, the choice of Pytorch 1.12 framework, the processor is intel i9-11900, the configuration of Nvidia GeForce RTX A4000 graphics card, the specific experimental environment hardware and software configuration as presented in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Experimental environment configuration</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Configuration name</th>
<th align="left">Configuration information</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Operating system</td>
<td align="left">Windows10</td>
</tr>
<tr>
<td align="left">Memory</td>
<td align="left">64G</td>
</tr>
<tr>
<td align="left">CPU</td>
<td align="left">Inter (R) Core (TM) i9-11900@2.50&#x2005;GHz 2.50&#x2005;GHz</td>
</tr>
<tr>
<td align="left">GPU</td>
<td align="left">Nvidia GeForce RTX A4000</td>
</tr>
<tr>
<td align="left">Language</td>
<td align="left">Python 3.7</td>
</tr>
<tr>
<td align="left">GPU acceleration</td>
<td align="left">Cuda11.6 Cudnn11.5</td>
</tr>
<tr>
<td align="left">Software environment</td>
<td align="left">Anaconda, PyCharm, Pytorch</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3"><label>4.3</label><title>Model Evaluation Indicators</title>
<p>To objectively and accurately judge the modified model performance, this paper adopts Precision (P), Recall (R), Average Precision (AP), and Mean Average Precision (mAP) as the hard metrics of the model; and adopts Frames Per Second (FPS), Parameters (Params) and model Size as an auxiliary indicator to measure the performance of the improved method. The hard indicators are calculated as follows.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mo>&#x222B;</mml:mo></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>continuous</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>discrete</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mi>K</mml:mi></mml:mfrac></mml:math></disp-formula></p>
<p>In the formula, <italic>TP</italic> (True Positives) means the count of positive samples which are correctly identified; <italic>FP</italic> (False Positives) means the count of positive samples detected incorrectly; <italic>FN</italic> (False Negatives) means the count of negative samples identified incorrectly; <italic>K</italic> means the total count of identified categories (7 in this paper); <italic>AP</italic>(i) means the AP value of the ith category; <italic>mAP</italic> is the average value of the detected APs of various types of surface litter. The larger of <italic>mAP</italic> value, the higher the model detection accuracy.</p>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Results and Analysis</title>
<sec id="s5_1"><label>5.1</label><title>Network Training</title>
<p>Before the network starts working, it needs to set some parameters of the network, so that the improved YOLOv5s network can reach the best training effect. This paper used the migration training idea to use the training weights of YOLOv5s over the COCO dataset as pretraining weights of the model backbone network and prevent random values of backbone weights to raise feature extraction. To further strengthen the generalization of the dataset, use the Mosaic&#x2009;&#x002B;&#x2009;Mixup data boosting technique at the input side, and the data boosting technique is performed in the first 70 rounds according to the probability of 50&#x0025; in each round. The input picture resolution is 480&#x2009;&#x00D7;&#x2009;480, and the model training batch size is 8. Meanwhile, the Adam optimizer with random gradient descent was selected to train the network optimally, setting the momentum to 0.937, the initial learning rate to 0.001, and the weight decay coefficient to 0 for 100 rounds of training.</p>
<p>To test the advantages of the modified method, the training process is visualized using the visualization tool Tensorboard, as shown in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. The horizontal axis denotes the count of training rounds while the vertical axis denotes the corresponding parameters. From the figure, the improved model has a faster and more stable accuracy increase in the initial stage, and the overall improvement effect is slightly superior than before, and the advanced model has smaller loss values and faster convergence of the loss function.</p>
<fig id="fig-11"><label>Figure 11</label><caption><title>Comparison of model performance visualization before and after improvement. (a) average precision value (mAP) (b) loss function (Loss)</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-11.tif"/></fig>
</sec>
<sec id="s5_2"><label>5.2</label><title>Ablation Experiments</title>
<p>To verify the usefulness of the various methods proposed in this paper, we did ablation experiments based on YOLOv5s using the same hyperparameters and training strategies, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. Experiment 1 is the original YOLOv5s, compared with Experiment 1, Experiment 2 replaces the Ghost module; Experiment 3 adds the CBAM module; Experiment 4 changes the CIoU loss function to the Focal-EIoU loss function; Experiment 5 combines the three improved methods. It can be seen from Experiment 2 that the model Size, GFLOPs, and Param are reduced by almost 50&#x0025; after replacing the module with Ghost, but the reduction of model parameters leads to a decrease in accuracy. The results of Experiment 3 and Experiment 4 show that the addition of the CBAM module and the improvement of the Focal-EIoU loss function have improved the accuracy of the algorithm. The final results of Experiment 5 show that the model incorporating the three improved methods has a 3.1&#x0025; improvement in mAP, a 39&#x0025; reduction in Size, a 44&#x0025; reduction in GFLOPs, and a 40&#x0025; reduction in Params compared to the original YOLOv5s model. Therefore, the comprehensive performance of the enhanced algorithm proposed in this paper is superior to the original YOLOv5s about surface litter detection, and it is also more appropriate for application in the computing devices of litter-cleaning vessels.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Ablation experiment</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Experiment</th>
<th align="left">Models</th>
<th align="left">G</th>
<th align="left">C</th>
<th align="left">E</th>
<th align="left">mAP_0.5</th>
<th align="left">Size/MB</th>
<th align="left">GFLOPs</th>
<th align="left">Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">YOLOv5s</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left">0.952</td>
<td align="left">14.4</td>
<td align="left">16.0</td>
<td align="left">7.039</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">YOLOv5s-G</td>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left"/>
<td align="left"/>
<td align="left">0.937</td>
<td align="left">7.8</td>
<td align="left">8.1</td>
<td align="left">3.692</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">YOLOv5s-C</td>
<td align="left"/>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left"/>
<td align="left">0.974</td>
<td align="left">15.4</td>
<td align="left">16.6</td>
<td align="left">7.553</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">YOLOv5s-E</td>
<td align="left"/>
<td align="left"/>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left">0.969</td>
<td align="left">14.4</td>
<td align="left">15.8</td>
<td align="left">7.029</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">YOLOv5s-GCE</td>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left"><bold>&#x221A;</bold></td>
<td align="left">0.983</td>
<td align="left">8.8</td>
<td align="left">8.9</td>
<td align="left">4.216</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3"><label>5.3</label><title>Comparison Experiments</title>
<p>In addition, based on the above parameter settings, the same datasets were used to train on the original YOLOv5s, YOLOv7-tiny, YOLOv3, SSD, and Faster R-CNN algorithm models, and the metric comparison of each model is listed in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Comparison of metrics of mainstream algorithm</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Models</th>
<th align="left">mAP_0.5</th>
<th align="left">P/&#x0025;</th>
<th align="left">R/&#x0025;</th>
<th align="left">Size/MB</th>
<th align="left">FPS</th>
<th align="left">GFLOPs</th>
<th align="left">Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">YOLOv3</td>
<td align="left">0.916</td>
<td align="left">91.25</td>
<td align="left">82.99</td>
<td align="left">235.2</td>
<td align="left">71</td>
<td align="left">87.4</td>
<td align="left">61.556</td>
</tr>
<tr>
<td align="left">SSD</td>
<td align="left">0.907</td>
<td align="left">92.04</td>
<td align="left">67.56</td>
<td align="left">93.7</td>
<td align="left">50</td>
<td align="left">155.1</td>
<td align="left">24.414</td>
</tr>
<tr>
<td align="left">Faster R-CNN</td>
<td align="left">0.889</td>
<td align="left">56.22</td>
<td align="left">95.72</td>
<td align="left">108.4</td>
<td align="left">16</td>
<td align="left">919.0</td>
<td align="left">28.337</td>
</tr>
<tr>
<td align="left">YOLOv5s</td>
<td align="left">0.952</td>
<td align="left">91.96</td>
<td align="left">92.19</td>
<td align="left">14.4</td>
<td align="left">108</td>
<td align="left">16.0</td>
<td align="left">7.039</td>
</tr>
<tr>
<td align="left">YOLOv7-tiny</td>
<td align="left">0.967</td>
<td align="left">92.62</td>
<td align="left">95.61</td>
<td align="left">11.7</td>
<td align="left">128</td>
<td align="left">13.1</td>
<td align="left">6.024</td>
</tr>
<tr>
<td align="left">YOLOv5s-GCE</td>
<td align="left">0.983</td>
<td align="left">97.70</td>
<td align="left">97.17</td>
<td align="left">8.8</td>
<td align="left">121</td>
<td align="left">8.9</td>
<td align="left">4.216</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From the comparison of the performance of each model in <xref ref-type="table" rid="table-4">Table 4</xref>, the improved algorithm in this paper performs better than other algorithms in aspects of mAP, precision, recall, and model size. mAP improved by 3.1&#x0025; over the original YOLOv5s algorithm, 1.6&#x0025; over YOLOv7-tiny, 6.7&#x0025; over YOLOv3&#x0025;, 9.4&#x0025; over Faster R-CNN, and 7.6&#x0025; over SSD; A 39&#x0025; reduction in model size compared to the original YOLOv5s and a 25&#x0025; reduction compared to the YOLOv7-tiny; FPS on the GPU is 13 higher than the original YOLOv5s and slightly lower than the YOLOv7-tiny. Params has much smaller than YOLOv3, Faster R-CNN, and SSD, and 30&#x0025; less than YOLOv7-tiny. By comparing the performance of these models, the improved model has higher detection accuracy while ensuring real-time detection and lightweight volume of the algorithm.</p>

</sec>
<sec id="s5_4"><label>5.4</label><title>Experiment Analysis</title>
<p>For verifying the actual detection effect of the revised algorithm, test validation is performed using test set images with complex backgrounds. The confidence threshold for the test set is set to 0.5, and the practical detection capacity of the algorithm before and after enhancement is presented in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>. From the detection comparisons in <xref ref-type="fig" rid="fig-12">Figs. 12a</xref> and <xref ref-type="fig" rid="fig-12">12b</xref>, the detection accuracy of the improved method is greater than before. The original algorithm is prone to miss-detection when detecting images with more complex backgrounds, especially small target categories, and obscured target categories, while the improved algorithm in this paper adopts the CBAM module and optimized loss function, which can detect targets more accurately and effectively decline the miss-detection rate while significantly enhancing the feature acquisition capability of the network. Therefore, the algorithm proposed by this paper has better detection accuracy and detection speed while maintaining a small size, which is both feasible and superior and can fulfill the practical demands of real-time detection.</p>
<fig id="fig-12"><label>Figure 12</label><caption><title>Comparison chart of model detection results. (a) original YOLOv5s test results (b) improved YOLOv5s detection results</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_39451-fig-12.tif"/></fig>
</sec>
</sec>
<sec id="s6"><label>6</label><title>Conclusion</title>
<p>For the traditional manual cleaning of water surface litter is time-consuming and inefficient, we proposed a lightweight water surface litter detection method by improved YOLOv5s to provide core technical support for water surface litter cleaning vessels. Based on the YOLOv5s network model, this model introduces a lightweight GhostNet network to replace the ordinary convolutional blocks and C3 modules in the original network, decreases the count of network parameters and computational overhead; embeds the CBAM module in the network to reinforce the network&#x2019;s capability to extract object feature information; optimize the loss function and use the Focal-EIoU loss function to enhance the regression accuracy and solve the positive and negative sample proportion imbalance problem. After practical testing, the mAP of the improved algorithm reaches 98.3&#x0025; and the model size is 8.8&#x2005;MB, which is 39&#x0025; less, with 44&#x0025; fewer GFLOPs, and 40&#x0025; fewer parameters compared to the original model. The single image test speed is about 8&#x2005;ms, which has high detection accuracy and detection speed while maintaining lightweight volume, it can fulfill the needs of real-time detection and processing of water surface litter.</p>
<p>After that, it will further enrich the count of categories in the water surface litter dataset, improve the generalization capability of the model, optimize the network structure, compress the model volume, facilitate the subsequent terminal deployment, and realize real-time detection in mobile.</p>
</sec>
</body>
<back>
<ack>
<p>The authors gratefully acknowledge the support of China University Industry-University Research Innovation Fund Project, Science and Technology Innovation R&#x0026;D Project of the State General Administration of Sports of China, Major Project of Philosophy and Social Science Research in Higher Education Institutions in Hubei Province, Key Project of Hubei Provincial Key Laboratory of Intelligent Transportation Technology and Equipment Open Fund.</p>
</ack>
<sec><title>Funding Statement</title>
<p>Support for this work was in part from the China University Industry-University Research Innovation Fund Project (No. 2022BL052), author B.T, <ext-link ext-link-type="uri" xlink:href="https://www.cutech.edu.cn">https://www.cutech.edu.cn</ext-link>; in part by the Science and Technology Innovation R&#x0026;D Project of the State General Administration of Sports of China (No. 22KJCX024), author B.T, <ext-link ext-link-type="uri" xlink:href="https://www.sport.gov.cn">https://www.sport.gov.cn</ext-link>; in part by the Major Project of Philosophy and Social Science Research in Higher Education Institutions in Hubei Province (No. 21ZD054), author B.T, <ext-link ext-link-type="uri" xlink:href="https://jyt.hubei.gov.cn">https://jyt.hubei.gov.cn</ext-link>; Key Project of Hubei Provincial Key Laboratory of Intelligent Transportation Technology and Equipment Open Fund (No. 2022XZ106), author B.T, <ext-link ext-link-type="uri" xlink:href="https://hbpu.edu.cn">https://hbpu.edu.cn</ext-link>.</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The manuscript is submitted without conflict of interest and all authors have approved the manuscript for public release. An original study of the work described has not been submitted or published elsewhere, the revised manuscript has been approved by all listed authors.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Tong</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>China is implementing &#x201C;Garbage classification&#x201D; action</article-title>,&#x201D; <source>Environmental Pollution</source>, vol. <volume>259</volume>, no. <issue>11307</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>2</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Kong</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Qiu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>IWSCR: An intelligent water surface cleaner robot for collecting floating garbage</article-title>,&#x201D; in <source>IEEE Transactions on Systems, Man, and Cybernetics: Systems</source>, vol. <volume>51</volume>, no. <issue>10</issue>, pp. <fpage>6358</fpage>&#x2013;<lpage>6368</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Gao</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Design and implementation of small waters intelligent garbage cleaning robot system based on raspberry pi</article-title>,&#x201D; <source>Science Technology and Engineering</source>, vol. <volume>19</volume>, no. <issue>34</issue>, pp. <fpage>239</fpage>&#x2013;<lpage>247</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Dilshad</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ullah</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Kim</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Seo</surname></string-name></person-group>, &#x201C;<article-title>LocateUAV: Unmanned aerial vehicle location estimation via contextual analysis in an IoT environment</article-title>,&#x201D; <source>Internet of Things Journal</source>, vol. <volume>10</volume>, no. <issue>5</issue>, pp. <fpage>4021</fpage>&#x2013;<lpage>4033</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Research on image target detection algorithm based on depth learning</article-title>,&#x201D; <source>Foreign Electronic Measurement Technology</source>, vol. <volume>39</volume>, no. <issue>8</issue>, pp. <fpage>34</fpage>&#x2013;<lpage>39</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Lite-hrnet: A lightweight high-resolution network</article-title>,&#x201D; in <conf-name>2021 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Nashville, TN, USA</conf-loc>, pp. <fpage>10440</fpage>&#x2013;<lpage>10450</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Su</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Taylor</surname></string-name></person-group>, &#x201C;<article-title>An elevator button recognition method combining yolov5 and ocr</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>75</volume>, no. <issue>1</issue>, pp. <fpage>117</fpage>&#x2013;<lpage>131</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Donahue</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Darrell</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Malik</surname></string-name></person-group>, &#x201C;<article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>,&#x201D; in <conf-name>2014 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Columbus, OH, USA</conf-loc>, pp. <fpage>580</fpage>&#x2013;<lpage>587</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name></person-group>, &#x201C;<article-title>Fast R-CNN</article-title>,&#x201D; in <conf-name>2015 IEEE Int. Conf. on Computer Vision (ICCV)</conf-name>, <conf-loc>Santiago, Chile</conf-loc>, pp. <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, <string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Faster R-CNN: Towards real-time object detection with region proposal networks</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>39</volume>, no. <issue>6</issue>, pp. <fpage>1137</fpage>&#x2013;<lpage>1149</lpage>, <year>2017</year>; <pub-id pub-id-type="pmid">27295650</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Redmon</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Divvala</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Farhadi</surname></string-name></person-group>, &#x201C;<article-title>You only look once: Unified, real-time object detection</article-title>,&#x201D; in <conf-name>2016 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Las Vegas, NV, USA</conf-loc>, pp. <fpage>779</fpage>&#x2013;<lpage>788</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Redmon</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Farhadi</surname></string-name></person-group>, &#x201C;<article-title>YOLO9000: Better, faster, stronger</article-title>,&#x201D; in <conf-name>2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, pp. <fpage>6517</fpage>&#x2013;<lpage>6525</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Anguelov</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Erhan</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Szegedy</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Reed</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>SSD: Single shot multibox detector</article-title>,&#x201D; in <conf-name>2016 European Conf. on Computer Vision (ECCV)</conf-name>, <conf-loc>Amsterdam, Netherlands</conf-loc>, pp. <fpage>21</fpage>&#x2013;<lpage>37</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y. -G.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Chen</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>DSOD: Learning deeply supervised object detectors from scratch</article-title>,&#x201D; in <conf-name>2017 IEEE Int. Conf. on Computer Vision (ICCV)</conf-name>, <conf-loc>Venice, Italy</conf-loc>, pp. <fpage>1937</fpage>&#x2013;<lpage>1945</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Mathias</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dhanalakshmi</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Narayanamoorthi</surname></string-name></person-group>, &#x201C;<article-title>Deep neural network driven automated underwater object detection</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>70</volume>, no. <issue>3</issue>, pp. <fpage>5251</fpage>&#x2013;<lpage>5267</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Woo</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Lee</surname></string-name> and <string-name><given-names>I. S.</given-names> <surname>Kweon</surname></string-name></person-group>, &#x201C;<article-title>CBAM: Convolutional block attention module</article-title>,&#x201D; in <conf-name>Proc. the European Conf. on Computer Vision (ECCV)</conf-name>, <conf-loc>Munich, Germany</conf-loc>, pp. <fpage>3</fpage>&#x2013;<lpage>19</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Jia</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Focal and efficient IoU loss for accurate bounding box regression</article-title>,&#x201D; <comment>arXiv preprint</comment>, pp. <fpage>146</fpage>&#x2013;<lpage>157</lpage>, <year>2021</year>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/2101.08158">https://arxiv.org/pdf/2101.08158</ext-link></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Ye</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Distance-IoU loss: Faster and better learning for bounding box regression</article-title>,&#x201D; <comment>arXiv preprint</comment>, pp. <fpage>12993</fpage>&#x2013;<lpage>13000</lpage>, <year>2019</year>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1911.08287v1">https://arxiv.org/abs/1911.08287v1</ext-link></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Matsumoto</surname></string-name></person-group>, &#x201C;<article-title>Ship image recognition using HOG</article-title>,&#x201D; <source>Journal of Japan Institute of Navigation</source>, vol. <volume>129</volume>, pp. <fpage>105</fpage>&#x2013;<lpage>112</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Kaido</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yamamoto</surname></string-name> and <string-name><given-names>T.</given-names> <surname>Hashimoto</surname></string-name></person-group>, &#x201C;<article-title>Examination of automatic detection and tracking of ships on camera image in marine environment</article-title>,&#x201D; in <conf-name>2016 IEEE Techno-Ocean (Techno-Ocean)</conf-name>, <conf-loc>Kobe, Japan</conf-loc>, pp. <fpage>58</fpage>&#x2013;<lpage>63</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Shen</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Real-time water surface object detection based on improved faster R-CNN</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>19</volume>, no. <issue>16</issue>, pp. <fpage>3523</fpage>&#x2013;<lpage>3539</lpage>, <year>2019</year>; <pub-id pub-id-type="pmid">31408971</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Panwar</surname></string-name>, <string-name><given-names>P. K.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>M. K.</given-names> <surname>Siddiqui</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Morales-Menendez</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Bhardwaj</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>AquaVision: Automating the detection of waste in water bodies using deep transfer learning</article-title>,&#x201D; <source>Case Studies in Chemical and Environmental Engineering</source>, vol. <volume>2</volume>, no. <issue>100026</issue>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Kong</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wu</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>A modified YOLOv3 detection method for vision-based water surface garbage capture robot</article-title>,&#x201D; <source>International Journal of Advanced Robotic Systems</source>, vol. <volume>17</volume>, no. <issue>3</issue>, pp. <fpage>1729</fpage>&#x2013;<lpage>8806</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Jocher</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Stoken</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Borovec</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Chaurasia</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Changyu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Ultralytics/yolov5: V5.0-YOLOv5-p6 models, AWS, supervisely and YouTube integrations</article-title>,&#x201D; <year>2021</year>. <pub-id pub-id-type="doi">10.5281/zenodo.4679653</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. -Y.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Doll&#x00E1;r</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name>, <string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Hariharan</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Feature pyramid networks for object detection</article-title>,&#x201D; in <conf-name>2017 IEEE Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, pp. <fpage>936</fpage>&#x2013;<lpage>944</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Qi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Qin</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Shi</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>Path aggregation network for instance segmentation</article-title>,&#x201D; in <conf-name>2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Salt Lake City, UT, USA</conf-loc>, pp. <fpage>8759</fpage>&#x2013;<lpage>8768</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Zhong</surname></string-name></person-group>, &#x201C;<article-title>Survey of deep learning model compression and acceleration</article-title>,&#x201D; <source>Journal of Software</source>, vol. <volume>32</volume>, no. <issue>1</issue>, pp. <fpage>68</fpage>&#x2013;<lpage>92</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Ge</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Shen</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Survey of lightweight netural network</article-title>,&#x201D; <source>Journal of Software</source>, vol. <volume>31</volume>, no. <issue>9</issue>, pp. <fpage>2627</fpage>&#x2013;<lpage>2653</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>F. N.</given-names> <surname>Iandola</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>M. W.</given-names> <surname>Moskewicz</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ashraf</surname></string-name>, <string-name><given-names>W. J.</given-names> <surname>Dally</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and &#x003C;0.5&#x2005;MB model size</article-title>,&#x201D; <comment>arXiv preprint</comment>, <year>2016</year>. <ext-link ext-link-type="uri" xlink:href="https://arvix.org/pdf/1602.07360">https://arvix.org/pdf/1602.07360</ext-link></mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Lin</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>ShuffleNet: An extremely efficient convolutional neural network for mobile devices</article-title>,&#x201D; in <conf-name>2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Salt Lake City, UT, USA</conf-loc>, pp. <fpage>6848</fpage>&#x2013;<lpage>6856</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>H. T.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zheng</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>ShuffleNet v2: Practical guidelines for efficient cnn architecture design</article-title>,&#x201D; in <conf-name>2018 European Conf. on Computer Vision (ECCV)</conf-name>, <conf-loc>Munich, Germany</conf-loc>, pp. <fpage>122</fpage>&#x2013;<lpage>138</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>A. G.</given-names> <surname>Howard</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Kalenichenko</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>MobileNets: Efficient convolutional neural networks for mobile vision applications</article-title>,&#x201D; <comment>arXiv preprint</comment>, <year>2017</year>. <ext-link ext-link-type="uri" xlink:href="https://arvix.org/abs/1704.04861">https://arvix.org/abs/1704.04861</ext-link></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sandler</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Howard</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Zhmoginov</surname></string-name> and <string-name><given-names>L. -C.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Mobilenetv2: Inverted residuals and linear bottlenecks</article-title>,&#x201D; in <conf-name>2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Salt Lake City, UT, USA</conf-loc>, pp. <fpage>4510</fpage>&#x2013;<lpage>4520</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Howard</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sandler</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Searching for mobilenetv3</article-title>,&#x201D; in <conf-name>2019 IEEE/CVF Int. Conf. on Computer Vision (ICCV)</conf-name>, <conf-loc>Seoul, Korea (South)</conf-loc>, pp. <fpage>1314</fpage>&#x2013;<lpage>1324</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="thesis"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>Research on scene image classification algorithm based on deep learning</article-title>,&#x201D; M. S. Dissertation, <publisher-name>Beijing University of Posts and Telecommunications</publisher-name>, <publisher-loc>China</publisher-loc>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>SNAS: Stochastic neural architecture search</article-title>,&#x201D; in <conf-name>2019 European Conf. on Computer Vision (ECCV)</conf-name>, <conf-loc>New Orleans, LA, USA</conf-loc>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Shen</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Squeeze-and-excitation networks</article-title>,&#x201D; in <conf-name>2018 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Salt Lake City, UT, USA</conf-loc>, pp. <fpage>7132</fpage>&#x2013;<lpage>7141</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Cisse</surname></string-name>, <string-name><given-names>Y. N.</given-names> <surname>Dauphin</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Lopez-Paz</surname></string-name></person-group>, &#x201C;<article-title>Mixup: Beyond empirical risk minimization</article-title>,&#x201D; <comment>arXiv preprint</comment>, <year>2017</year>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1710.09412">https://arxiv.org/abs/1710.09412</ext-link></mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Xu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>GhostNet: More features from cheap operations</article-title>,&#x201D; in <conf-name>2020 IEEE/CVF Conf. on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Seattle, WA, USA</conf-loc>, 2020, pp. <fpage>1577</fpage>&#x2013;<lpage>1586</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Ke</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Liao</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>Research on hierarchical decomposition of convolution neural network</article-title>,&#x201D; <source>Computer Engineering</source>, vol. <volume>45</volume>, no. <issue>11</issue>, pp. <fpage>191</fpage>&#x2013;<lpage>197</lpage>, <year>2019</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>