<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79851</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.079851</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Accurate Real-Time Measurement of Small and Irregular Road Abandoned Objects Using a Lightweight Vision-Based Framework</article-title>
<alt-title alt-title-type="left-running-head">Accurate Real-Time Measurement of Small and Irregular Road Abandoned Objects Using a Lightweight Vision-Based Framework</alt-title>
<alt-title alt-title-type="right-running-head">Accurate Real-Time Measurement of Small and Irregular Road Abandoned Objects Using a Lightweight Vision-Based Framework</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Tang</surname><given-names>Ying</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Ma</surname><given-names>Chuanyi</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Guo</surname><given-names>Feng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>fengg@sdu.edu.cn</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Sun</surname><given-names>Wenhao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Qilu Transportation, Shandong University</institution>, <addr-line>Jinan</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Shandong Hi-Speed Group Co., Ltd.</institution>, <addr-line>Jinan</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Feng Guo. Email: <email>fengg@sdu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>64</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_79851.pdf"></self-uri>
<abstract>
<p>Road Abandoned Objects (RAOs) pose significant threats to traffic safety, particularly due to their small size, irregular shapes, and unpredictable distribution in complex road environments. The primary objective of this study is to develop an accurate and real-time detection framework for RAOs while maintaining low computational cost for practical deployment. To achieve this, we propose RAO-YOLO, a lightweight vision-based detection framework built upon an enhanced YOLO architecture. Specifically, a Mixed Aggregation Network (MANet) is introduced to improve multi-scale feature representation, and a Lightweight Shared Detail-Enhanced Detection (LSDD) head is designed to enhance localization accuracy for small and irregular objects. Furthermore, a Focal-MPDIoU loss function is proposed to address sample imbalance and geometric irregularity during training. Extensive experiments conducted on the RAOD dataset demonstrate that the proposed method achieves superior performance compared to state-of-the-art detectors, achieving a mAP@0.5:0.95 of 56.1% while maintaining real-time inference speed. These results validate the effectiveness of the proposed framework for practical intelligent transportation applications.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Road abandoned objects</kwd>
<kwd>real-time object detection</kwd>
<kwd>road safety</kwd>
<kwd>intelligent transportation systems</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Natural Science Foundation of China</funding-source>
<award-id>52308457</award-id>
</award-group>
<award-group id="awg2">
<funding-source>China Postdoctoral Science Foundation</funding-source>
<award-id>2024M761811</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Natural Science Foundation of Shandong Province</funding-source>
<award-id>ZR2023QE220</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Traffic incidents pose serious safety and economic challenges worldwide, resulting in both substantial financial losses and travel disruptions [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]. Among the various causes, road abandoned objects (RAOs), referring to unexpected static obstacles such as debris, cargo, or fragments left on the roadway, represent a notable safety hazard.</p>
<p>In this study, real-time detection refers to achieving inference speeds that satisfy practical deployment requirements (typically above 30 FPS) while maintaining reliable detection accuracy. However, detecting small and irregular RAOs in real-time remains a challenging task due to their diverse shapes, small scales, and complex roadside environments. To address this issue, current detection practices primarily rely on static roadside sensors. While these systems can identify anomalies by monitoring traffic flow, they suffer from high false alarm rates and poor adaptability to complex road environments. Thus, there is an urgent need for accurate and efficient RAO detection to ensure smooth operations and maintain traffic safety. An illustration of RAOs highlighted with red boxes, is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Illustration of road abandoned objects (RAOs). Red bounding boxes highlight the RAO targets in the scene.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-1.tif"/>
</fig>
<p>Traditional sensor-based and rule-based approaches have been widely used for RAO detection. Sensor-based methods rely on hardware devices such as inductive loops, acoustic sensors, and passive infrared detectors [<xref ref-type="bibr" rid="ref-3">3</xref>], while rule-based approaches depend on handcrafted thresholds, background modeling, and heuristic decision rules to identify anomalies. These systems identify anomalies by monitoring traffic flow patterns; however, on the one hand, their fixed-location nature restricts spatial coverage, and on the other hand, their adaptability to complex and dynamic traffic conditions is limited [<xref ref-type="bibr" rid="ref-4">4</xref>]. As a result, detection accuracy can be significantly degraded under varying lighting, weather, and occlusion conditions. Recent advances in computer vision provide solutions for overcoming these limitations. With traffic cameras now widely deployed, modern object detection algorithms indicate strong capabilities in recognizing multi-scale targets under real-time constraints, offering a promising foundation for RAO detection on highways [<xref ref-type="bibr" rid="ref-5">5</xref>]. These limitations highlight the need for more adaptive vision-based detection approaches capable of handling complex roadside scenarios.</p>
<p>Early monocular vision-based RAO detection methods relied on edge extraction and background modeling. For instance, the Gaussian Mixture Model (GMM) separates foreground objects by modeling pixel intensities [<xref ref-type="bibr" rid="ref-6">6</xref>], while the Active Contour Model (ACM) detects boundaries via energy minimization [<xref ref-type="bibr" rid="ref-7">7</xref>]. Wavelet filtering and Bayesian decision models have also shown promise in controlled settings [<xref ref-type="bibr" rid="ref-8">8</xref>]. However, these approaches are highly sensitive to lighting, weather, and occlusions, limiting their real-world robustness. To improve geometric awareness, stereo vision techniques were introduced. Using calibrated camera pairs, methods like V-disparity [<xref ref-type="bibr" rid="ref-9">9</xref>] and adaptive baseline stereo [<xref ref-type="bibr" rid="ref-10">10</xref>] estimate object height and distance from disparity maps. Some systems reconstruct 3D point clouds for obstacle clustering and classification [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Although more robust to illumination changes, stereo setups require precise calibration, lack scalability, and often need manual tuning&#x2014;hindering large-scale roadside deployment.</p>
<p>Recent advances in AI, particularly convolutional neural networks (CNNs), offer a more adaptable solution. CNNs have achieved remarkable success in object detection across diverse scenarios. Models like Faster R-CNN [<xref ref-type="bibr" rid="ref-13">13</xref>], SSD [<xref ref-type="bibr" rid="ref-14">14</xref>], and YOLO [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>] enable accurate, real-time object detection and are well-suited for video surveillance. Yet, generic detectors struggle with small, irregular, or context-dependent RAOs. Their practical use remains limited by the scarcity of large-scale RAO-specific datasets and challenges in seamless real-time integration. In particular, standard detection architectures often suffer from insufficient feature representation for small targets and limited localization accuracy for irregularly shaped objects.</p>
<p>Despite recent progress in object detection, effectively detecting small and irregular RAOs under strict real-time constraints remains challenging, we propose an enhanced detection framework, termed RAO-YOLO, built upon YOLOv11m and incorporating targeted improvements. Specifically, our main contributions are as follows:<list list-type="bullet">
<list-item>
<p>We adopt a Mixed Aggregation Network (MANet) to replace the original YOLO feature extraction module. By aggregating multi-scale and multi-branch features, this module significantly enhances the network&#x2019;s representation ability for complex roadside objects.</p></list-item>
<list-item>
<p>We propose a lightweight and efficient detection head, termed LSDD (Lightweight Shared Detail-Enhanced Detection), which improves both classification and localization accuracy through shared detail-enhancing layers and parallel prediction branches.</p></list-item>
<list-item>
<p>We introduce the loss function Focal-MPDIoU, which integrates a Focal IoU mapping and corner distance penalty to better fit the characteristics of small and irregularly shaped RAOs, leading to improved training stability and performance.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>This section reviews existing research relevant to RAO detection. We first present representative datasets and benchmarks, as they provide the foundation for method development and performance evaluation in this domain. We then summarize traditional vision-based methods and AI-based approaches, highlighting their strengths and limitations in complex traffic environments.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Datasets and Benchmarks</title>
<p>Although general-purpose datasets (e.g., Cityscapes [<xref ref-type="bibr" rid="ref-17">17</xref>], KITTI [<xref ref-type="bibr" rid="ref-18">18</xref>]) support road scene understanding, they lack annotations for abandoned objects, limiting their applicability to RAO detection.</p>
<p>To fill this gap, specialized datasets have emerged. LostAndFound [<xref ref-type="bibr" rid="ref-19">19</xref>] provides 2102 annotated frames of small obstacles from urban roads but suffers from limited scale and diversity. CAOS [<xref ref-type="bibr" rid="ref-20">20</xref>], based on BDD100K [<xref ref-type="bibr" rid="ref-21">21</xref>], targets two anomaly categories, while RoadObstacle21 [<xref ref-type="bibr" rid="ref-22">22</xref>] offers 321 high-resolution, pixel-annotated images of road-surface obstacles, yet both lack scene variability for robust highway generalization. The recently released RAOD dataset [<xref ref-type="bibr" rid="ref-23">23</xref>] addresses these shortcomings with a large-scale, diverse collection of real-world traffic videos featuring various RAOs under diverse conditions, establishing a realistic and challenging benchmark for practical RAO detection in both urban and highway settings.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Traditional Abandoned Object Detection Methods</title>
<p>Early approaches to abandoned object detection relied on handcrafted features and heuristic rules. Motion-based methods (e.g., frame differencing, optical flow [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]) identify static anomalies against moving traffic but fail under camera motion or when RAOs appear in sparse traffic. Background subtraction techniques including GMM, median filtering, and codebook models are efficient but highly vulnerable to shadows, reflections, rain, and illumination changes, leading to excessive false alarms.</p>
<p>Background subtraction methods, including Gaussian Mixture Models (GMM), temporal median filtering, and codebook models, formed the backbone of many early systems [<xref ref-type="bibr" rid="ref-25">25</xref>]. By learning a scene&#x2019;s background over time, they flag foreground anomalies. Although computationally efficient and suitable for real-time use, their accuracy is easily compromised by shadows, reflections, rain, or low-light conditions, leading to false positives or missed detections. Post-processing often employs appearance-based features like color histograms, edge maps, HOG [<xref ref-type="bibr" rid="ref-26">26</xref>], and LBP [<xref ref-type="bibr" rid="ref-27">27</xref>] to refine results. Yet these features lack robustness to intra-class variation and cannot generalize to unseen object categories, limiting their effectiveness in open-set, real-world traffic scenarios.</p>
<p>Moreover, most traditional methods depend on fixed thresholds and static rules, making them inflexible to varying lighting, camera angles, road layouts, or seasonal changes. While efficient and easy to deploy, their reliance on handcrafted heuristics and sensitivity to environmental dynamics severely constrain generalization across diverse highway and urban settings.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>AI-Based Detection Methods</title>
<p>Deep learning has significantly advanced RAO detection by enabling automatic learning of discriminative features. However, RAO detection differs from general object detection due to the small scale, irregular shapes, and ambiguous appearance of targets, which pose additional challenges for existing models. CNN-based detectors like Faster R-CNN [<xref ref-type="bibr" rid="ref-13">13</xref>], SSD [<xref ref-type="bibr" rid="ref-14">14</xref>], and YOLO [<xref ref-type="bibr" rid="ref-15">15</xref>] are widely used for their ability to learn hierarchical representations from annotated data, showing strong generalization&#x2014;especially when trained on large-scale datasets such as SYNTHIA [<xref ref-type="bibr" rid="ref-28">28</xref>]. However, these models are primarily designed for regular and medium-to-large objects, and often struggle to preserve discriminative features for small and irregular RAOs, leading to missed detections or inaccurate localization. Bakirci [<xref ref-type="bibr" rid="ref-29">29</xref>] comprehensively evaluate YOLOv8 variants in aerial traffic monitoring, offering insights into their robustness under varying illumination, density, and occlusion. Some works further integrate semantic segmentation to better distinguish RAOs from background clutter.</p>
<p>Transformer-based models like DETR and its variants [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>] capture global context, improving performance in occluded or low-contrast scenes. Streamlined architectures such as D-FINE [<xref ref-type="bibr" rid="ref-32">32</xref>] and RT-DETR [<xref ref-type="bibr" rid="ref-33">33</xref>] enable high-precision, real-time detection, their high computational cost and latency make them less suitable for real-time RAO detection in roadside deployment scenarios.</p>
<p>To mitigate sensor limitations under adverse conditions, recent studies explore multimodal fusion (RGB, LiDAR, radar) via feature- or decision-level strategies [<xref ref-type="bibr" rid="ref-34">34</xref>], enhancing robustness for small or low-contrast RAOs. However, these approaches face challenges in cost, complexity, latency, and computational overhead, hindering real-time highway deployment. Lightweight designs (e.g., MelNet [<xref ref-type="bibr" rid="ref-35">35</xref>], TinyDet [<xref ref-type="bibr" rid="ref-36">36</xref>]) and compression techniques (e.g., pruning, quantization) aim to address these issues. Despite progress, achieving both high accuracy and efficiency in dynamic, real-world traffic environments remains an open challenge.</p>
<p>Despite the progress of existing methods, they remain insufficient for RAO detection due to several key limitations, including inadequate feature representation for small-scale objects that often leads to missed detections, poor modeling of irregular object geometry resulting in inaccurate localization, and the inherent difficulty in balancing detection accuracy with computational efficiency under real-time constraints. To address these challenges, we propose RAO-YOLO, which enhances multi-scale feature representation, improves localization for irregular objects, and maintains real-time performance.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>To address the aforementioned challenges in RAO detection, including insufficient feature representation for small-scale objects, inaccurate localization of irregular targets, and the difficulty of maintaining real-time performance, we propose an improved detection framework termed RAO-YOLO. Built upon the baseline YOLO architecture, the proposed method introduces several targeted enhancements, where each component is explicitly designed to tackle a specific limitation of existing approaches. (1) Mixed Aggregation Network (MANet) module that enhances multi-scale feature aggregation and strengthens the representation of small objects; (2) Lightweight Shared Detail-Enhanced Detection (LSDD) head that improves localization accuracy for irregular targets while maintaining low computational overhead; (3) Focal-MPDIoU loss function designed to address sample imbalance and irregular object shapes during training.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Overview of Model Architecture</title>
<p>RAO-YOLO builds on YOLOv11 to tackle RAO detection in complex highway scenes. While preserving the one-stage paradigm, it modifies both backbone and detection head (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>). The backbone retains original C3K2 modules in shallow layers for efficient low-level feature extraction, while deeper stages replace C3K2 with MANet which is a module that fuses multi-scale features and attention mechanisms [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>] to better capture discriminative cues from small, irregular, or occluded RAOs. This design is particularly beneficial for RAO detection because abandoned objects in road scenes often occupy only a small number of pixels, exhibit irregular boundaries, and appear in cluttered backgrounds. Multi-scale aggregation helps preserve weak responses from small targets across different feature levels, while the attention mechanism suppresses redundant background interference and highlights informative regions. As a result, MANet improves both local detail sensitivity and global context awareness, which are essential for reliable detection in complex highway environments.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>RAO-YOLO model architecture. The proposed framework integrates MANet in the backbone and LSDD in the detection head to enhance feature representation and localization accuracy for small and irregular RAOs.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-2.tif"/>
</fig>
<p>The detection head is replaced with the lightweight LSDD head, which uses shared convolutions and detail enhancement to reduce complexity while improving fine-grained localization and classification especially effective for RAOs of varying scales and orientations. Combined with a tailored Focal-MPDIoU loss, RAO-YOLO achieves high accuracy and fast inference, making it suitable for real-time highway monitoring in intelligent transportation systems.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>MANet Module</title>
<p>In our proposed framework, we adapt the MANet module (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>), originally introduced in Hyper-YOLO [<xref ref-type="bibr" rid="ref-39">39</xref>], to address RAO detection challenges. While designed for general object detection, we integrate MANet into the deeper stages of the YOLOv11 backbone and tailor it to better capture small, irregular, and partially occluded roadside objects. The module retains its multi-branch structure (standard convolution, depthwise separable convolution, and bottleneck enhancement), but is re-parameterized to balance accuracy and efficiency for high-resolution highway imagery. Its lightweight attention mechanism further enhances fine-grained feature extraction in cluttered scenes. Compared with the conventional C3K2 module, the adapted MANet provides richer multi-scale representations with lower computational cost, improving robustness to scale variation and complex environments.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>MANet model structure.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-3.tif"/>
</fig>
<p>These architectural enhancements enable MANet to effectively capture both fine-grained local textures and high-level semantic cues with minimal redundancy. To explain the internal operation of MANet more intuitively, the module can be divided into: initial channel expansion, branch-wise feature transformation, and final feature fusion. First, the input feature is projected to a higher-dimensional intermediate representation to increase feature diversity. Second, the expanded feature is processed by different branches, where each branch focuses on complementary information. The core principle of the MANet module can be expressed as:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mover><mml:mo stretchy="false">&#x27F6;</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mover><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>After the input feature <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is processed by the first convolutional layer, the number of output <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> channels becomes twice the original (2C).
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mover><mml:mo stretchy="false">&#x27F6;</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mover><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mover><mml:mo stretchy="false">&#x27F6;</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mover><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mover><mml:mo stretchy="false">&#x27F6;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:mover><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>This decomposition allows different branches to focus on complementary information (e.g., compact features, enhanced textures, and preserved representations). Then the transformed feature F is divided into four different branches:</p>
<p><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> further compresses the channel through 1 &#x00D7; 1 convolution, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is extracted through a complex convolutional sequence that includes depthwise convolutions, as shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> features are split directly into two parts along the channel dimension while preserving the original information. The feature enhancement module uses <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>n</mml:mi></mml:math></inline-formula> Bottleneck modules to enhance the last feature. Finally, output result concatenates all the features and fuses them through 1 &#x00D7; 1 convolution, enabling effective integration of multi-scale and multi-type representations.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mi>B</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>k</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mi>e</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Assuming the input feature map has a dimension of <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:math></inline-formula>, the MANet module performs feature aggregation while maintaining the spatial resolution <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula> and adaptively enhancing informative channel responses. Compared with directly stacking standard convolutional blocks, this design provides a more effective balance between feature richness and computational cost. We adopt MANet because RAO detection requires sensitivity to subtle local cues while remaining efficient enough for real-time deployment. By incorporating this advanced module from Hyper-YOLO, RAO-YOLO achieves a favorable trade-off between detection accuracy and computational efficiency, significantly enhancing the model&#x2019;s practicality for real-world applications. Furthermore, the modular design of MANet allows for easy integration into various backbone architectures, making it a versatile component for modern object detection frameworks.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Optimized Detection Head</title>
<p>LSDD is a novel detection head module introduced in RAO-YOLO, specifically tailored for lightweight object detection tasks. The main idea of LSDD is to first align features at different scales and then enhance structural details using shared lightweight operators. This allows the detection head to focus more effectively on edge, texture, and directional information that is important for irregular RAOs. The structure of the proposed LSDD head is illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>LSDD head structure.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-4.tif"/>
</fig>
<p>Architecturally, LSDD adopts a multi-scale feature processing strategy, where dedicated 1 &#x00D7; 1 convolutional layers with group normalization [<xref ref-type="bibr" rid="ref-40">40</xref>] are applied to feature maps at different scales (i.e., P3, P4, and P5). These are followed by a shared detail-enhanced convolutional block, implemented as two consecutive DEConv_GN modules, which are applied across all scales. Group Normalization can better adapt to small-batch training and distributed training, enhancing the stability and robustness of the model.</p>
<p>This parameter-sharing mechanism significantly reduces the overall parameter count and computational burden, while maintaining strong feature extraction capabilities. The central component, DEConv, incorporates five distinct types of differential convolutions: center difference, horizontal difference, vertical difference, diagonal difference, and standard convolution. By capturing edge, texture, and directional features more effectively, these differential convolutions substantially enhance the model&#x2019;s ability to localize and classify objects, especially in challenging scenarios with small or densely packed targets. The calculation formula is as follows:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>D</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>D</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>D</mml:mi><mml:mi>E</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>5</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Among them, <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>G</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> stands for Group Normalization. <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents center, horizontal, vertical, diagonal difference and standard convolution respectively, and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the learnable weight.</p>
<p>From a technical perspective, LSDD leverages parameter sharing and detail enhancement to greatly improve computational efficiency and detection precision. The use of GN ensures robust training stability, even with small batch sizes or distributed training environments. Furthermore, the design of the regression and classification branches where the bounding box regression branch employs a learnable scale module and distribution focal loss (DFL), and the classification branch utilizes a dedicated 1 &#x00D7; 1 convolution enables the model to better handle multi-scale objects and complex backgrounds. This makes LSDD particularly suitable for real-time RAOs detection, where the detection of small, irregular, or partially occluded objects in complex road environments is critical for traffic safety and intelligent transportation systems. The lightweight and efficient design of LSDD allows for deployment of edge devices and embedded systems, meeting the stringent requirements of real-time RAOs detection in practical applications. The calculation formulas for the regression branch, classification branch, decoding and loss are as follows:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mover><mml:mi>b</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>j</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>B</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Among them, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> respectively represent the outputs of bounding box regression and category prediction branches, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the probability after softmax, and <italic>M</italic> &#x2212; 1 is the number of buckets divided. After decoding, the final output is <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>Y</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> indicates sigmoid activation, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> stands for location loss, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> stands for classification loss, <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the corresponding loss weights.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Improved Loss Function</title>
<p>YOLOv11 originally uses Complete IoU (CIoU) [<xref ref-type="bibr" rid="ref-41">41</xref>] as its regression loss, which improves localization by incorporating center distance and aspect ratio constraints. While effective for regular-shaped RAOs (e.g., cones, barriers), CIoU&#x2019;s rigid aspect ratio term struggles with irregular objects like fragments, spills, or scattered debris leading to localization errors for elongated or flattened targets.</p>
<p>Several IoU variants aim to address this: Shape-IoU enhances shape adaptability but suffers from high computational cost, limiting real-time use; Inner-IoU improves overlap quality for small/occluded objects but underperforms on large or highly irregular RAOs.</p>
<p>In object detection tasks, the training objective typically consists of multiple components, including classification and localization losses. Therefore, the overall optimization objective of the proposed RAO-YOLO detector can be formulated as a composite loss:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the classification loss used to predict object categories, and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the bounding box regression loss for object localization. The coefficients <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are weighting parameters used to balance the contributions of classification and localization during training.</p>
<p>To overcome these issues, we propose Focal-MPDIoU, a hybrid loss combining Focal-IoU [<xref ref-type="bibr" rid="ref-42">42</xref>] and Minimum Point Distance IoU (MPDIoU) [<xref ref-type="bibr" rid="ref-43">43</xref>]. MPDIoU eliminates aspect ratio dependence by measuring the minimum point distance between boxes, enabling accurate localization of irregular shapes. Focal-IoU further introduces a piecewise linear weighting scheme (with thresholds <italic>&#x03B1;</italic> and <italic>&#x03B2;</italic>) to suppress easy samples and emphasize hard cases such as small, occluded, or scale-varying RAOs. This design improves convergence stability, maintains high localization accuracy across diverse geometries, and supports robust real-time detection. The formulation of Focal-IoU is as follows:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Among them, <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> are hyperparameters defined over the interval (0, 1), and are typically set to 0 and 0.95, respectively. Here, <italic>&#x03B1;</italic> controls the balance between positive and negative samples, while <italic>&#x03B2;</italic> adjusts the weighting of hard samples, enabling the model to focus more on difficult instances during training. Meanwhile, MPDIoU improves localization accuracy by introducing a geometric penalty based on the Euclidean distance between the top-left and bottom-right corners of the predicted and ground-truth bounding boxes. Unlike traditional overlap-based metrics, MPDIoU captures subtle misalignments and is more robust to variations in aspect ratio which is an essential property when detecting elongated spills, fragmented RAOs, or non-rectangular roadside objects. The MPDIoU calculation formula is as follows:<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In this formula, <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>w</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>h</mml:mi></mml:math></inline-formula> are the image width and height. <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the predicted box&#x2019;s top-left and bottom-right coordinates, while <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are those of the ground-truth box. The Euclidean distances between corresponding corners are used to quantify localization error in MPDIoU.</p>
<p>By combining the adaptive sample weighting of Focal-IoU with the geometric sensitivity of MPDIoU, Focal-MPDIoU provides a more discriminative and shape-aware supervision signal, enabling RAO-YOLO to produce tighter, more consistent bounding boxes for challenging objects in real-world highway environments. The Focal-MPDIoU loss function calculation formula as follows:<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B2;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>M</mml:mi><mml:mi>P</mml:mi><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Compared to the other three loss functions, Focal-MPDIoU achieves a more balanced computational efficiency, ensuring detection accuracy while meeting real-time requirements. It provides more precise boundary localization for small targets, better handling of overlapping RAOs scenarios, and superior detection performance for partially occluding RAOs through its refined distance-based measurement approach (see <xref ref-type="fig" rid="fig-5">Fig. 5</xref>).</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison of IoU and Focal-MPDIoU penalty. (<bold>a</bold>) IoU overlap: The green box <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the ground truth, the red box <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the prediction, and the gray area is their intersection. The black dashed box C is the minimum enclosing. (<bold>b</bold>) Focal-MPDIoU penalty: In addition to the overlap, Focal-MPDIoU introduces penalties for the distances between the corresponding corners (<inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>) of the two boxes, as shown by the dashed lines. This helps to penalize misalignment and improve localization accuracy.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>RAOD Dataset</title>
<p>RAOD [<xref ref-type="bibr" rid="ref-23">23</xref>] is a large-scale, real-world benchmark specifically designed for abandoned object detection in road video surveillance which is a critical safety task for intelligent transportation systems. Addressing the limitations of autonomous-driving-centric datasets, RAOD offers greater diversity and realism through extensive CCTV footage captured across varied times, distances, and highway scenes. It includes 557 video sequences (over 500 with RAOs) and 18,891 pixel-annotated images, making it the largest open-source dataset for this task. Following the official dataset split provided by RAOD, the dataset is divided into 18,915 training images and 2600 testing images, which include both annotated RAO samples and additional background frames without abandoned objects. The data spans more than 70 highway locations and defines 10 common RAO categories, grouped into recognizable items (e.g., plastic bags, boxes, auto parts) and unidentifiable forms (e.g., barrel-, tabular-, or irregular-shaped objects). Representative statistics are shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Proportional distribution of object categories in the RAOD dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-6.tif"/>
</fig>
<p>To simplify the problem and enable models to focus on learning the common features of abandoned objects, all abandoned items are unified into a single &#x201C;abandoned&#x201D; class, thereby avoiding potential issues related to class imbalance that fine-grained categorization might introduce. This design allows the model to emphasize shared visual characteristics across categories, improving robustness in detecting small and irregular objects. It facilitates comprehensive evaluation experiments with various baseline models from diverse fields, including image object detection methods (e.g., YOLO series and DETR-based models) as well as representative approaches for road anomaly detection, which are more suitable for bounding-box-based evaluation in this task. Since the RAO detection task is formulated as a bounding-box-based detection problem, segmentation-based methods are not included in the comparison. <xref ref-type="table" rid="table-1">Table 1</xref> indicates the compared results between different RAOs datasets.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of publicly available datasets for road abandoned objects in the number of images, positive videos and types.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Images</th>
<th>Positives</th>
<th>Types</th>
</tr>
</thead>
<tbody>
<tr>
<td>LostAndFound [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>2104</td>
<td>112</td>
<td>9</td>
</tr>
<tr>
<td>RoadObstacle21 [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>327</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>RAOD [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>18,891</td>
<td>502</td>
<td>10</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Detail</title>
<p>All experiments, including those comparing RAO-YOLO with other object detection algorithms, are implemented using PyTorch and conducted on a consistent software and hardware platform to ensure fairness and reproducibility. The hardware setup consists of an NVIDIA RTX 3090 GPU with 24 GB of VRAM, an AMD EPYC 7642 48-core CPU, and 256 GB of system memory. The system operates on the Linux distribution, Ubuntu 22.04. The software stack used for training and evaluation includes Python, the PyTorch deep learning framework, and the NVIDIA CUDA Toolkit, among other supporting libraries. Following the official dataset split, the dataset consists of 18,915 training images and 2600 testing images, where the training set is used for model learning and the test set is used for final performance evaluation, while data augmentation is applied during training to reduce overfitting. The model was trained using the stochastic gradient descent (SGD) optimizer with an initial learning rate of 0.01 following the default training configuration of the framework, and no warm-up strategy was applied during training.</p>
<p>Regarding the experimental setup and hyperparameter selection, RAO-YOLO and other object detection models primarily rely on default configurations, with the exception of aggregated performance metrics, which are closely tracked throughout the training process. To ensure consistency and minimize the impact of input image resolution on the evaluation outcomes, all experiments were conducted using a standardized input size of 640 &#x00D7; 640 pixels, which is widely adopted in related research. For object detection algorithms, particularly those in the YOLO series, a unified performance metric that is commonly referred to as the <italic>fitness function</italic> is employed during training to guide model evaluation and selection. This function integrates multiple metrics into a single scalar score, with the default formulation defined as:<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>F</mml:mi><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mn>0.5</mml:mn><mml:mo>+</mml:mo><mml:mn>0.9</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mn>0.5</mml:mn><mml:mo>&#x003A;</mml:mo><mml:mn>0.95</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In this context, <italic>F</italic> denotes the fitness score, while mean Average Precision (mAP) quantifies the area under the precision&#x2013;recall (PR) curve. The key distinction between mAP@0.5 and mAP@0.5:0.95 lies in their respective IoU threshold criteria for bounding box matching. Based on the fitness function formulation, model selection is predominantly guided by performance on the mAP@0.5:0.95 metric. However, in specific tasks such as RAOs detection, researchers often emphasize precision and recall over mAP, due to the critical importance of reducing false positives and false negatives in intelligent transportation applications. The following sections provide a more detailed explanation of the evaluation metrics employed in this study.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metrics</title>
<p>For object detection tasks, a wide range of evaluation metrics [<xref ref-type="bibr" rid="ref-44">44</xref>] are commonly used. In the context of RAOs detection, we highlighted several widely adopted indicators, including precision, recall, and F1 score. Understanding these metrics requires familiarity with the foundational classification outcomes: true positive (TP), false positive (FP), true negative (TN), and false negative (FN). A true positive (TP) occurs when the model correctly identifies the presence of RAOs in an image. FP refers to an incorrect prediction where the model detects RAOs in an image that does not contain one. TN is when the model correctly predicts the absence of RAOs. FN arises when the model fails to detect RAOs that is actually present in the image. Based on these four quantities, the commonly used evaluation metrics are defined as follows:<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><xref ref-type="disp-formula" rid="eqn-25">Eqs. (25)</xref> and <xref ref-type="disp-formula" rid="eqn-26">(26)</xref> present precision (fraction of correct detections among all predictions), recall (fraction of detected RAOs out of all ground-truth instances), and their harmonic mean, the F1 score. However, for RAO detection&#x2014;where both identification and precise localization are critical&#x2014;single metrics like recall are insufficient. We therefore adopt mAP@0.5 (at IoU &#x003D; 0.5) and mAP@0.5:0.95 (averaged over IoU thresholds from 0.5 to 0.95 in 0.05 steps) as primary metrics. The former assesses object recognition capability, while the latter provides a stricter, more comprehensive evaluation of localization accuracy and robustness.</p>
<p>As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the confusion matrix highlights RAO-YOLO&#x2019;s strong performance: it achieves a high true positive count, indicating reliable RAO detection across diverse road conditions, while maintaining low false positives, demonstrating high precision and resistance to over-detection, crucial for minimizing false alarms. Compared to YOLOv11m, RAO-YOLO significantly increases true positives and reduces both false negatives and false positives, confirming its superior overall detection capability. The confusion matrix also reveals the distribution of misclassification cases, providing insight into the typical error patterns of the proposed model.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix of different model on the RAOD test set. The matrix illustrates the distribution of true positives, false positives, and false negatives, highlighting the improved detection performance of RAO-YOLO.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-7.tif"/>
</fig>
<p>While some false negatives still occur, true RAOs missed by the model, the overall pattern suggests that RAO-YOLO prioritizes precision without heavily compromising recall. This trade-off reflects the model&#x2019;s well-calibrated balance between accuracy and efficiency, making it particularly suitable for real-time roadside safety applications. From the confusion matrix, we can find it underscores RAO-YOLO&#x2019;s effectiveness in detecting varied and often subtle RAOs instances, confirming its advantage in both detection reliability and practical deployability compared to conventional models.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Comparative Experiments</title>
<p>We select YOLOv11m as the baseline due to its strong balance of accuracy, model complexity, and inference speed. While numerous detectors exist, YOLOv11m consistently excels due to architectural enhancements that support both efficiency and precision, crucial for real-time industrial systems. For a fair comparison, we evaluate it against recent models from the YOLOm and YOLOx series, as well as transformer-based RT-DETR, covering different design paradigms (CNN vs. transformer) and accuracy-efficiency trade-offs. All models are tested on the RAOD dataset under identical training and hardware conditions. Since the proposed MANet module operates on multi-scale feature maps without significantly increasing channel dimensions, the overall computational complexity remains comparable to the baseline YOLO architecture while improving feature representation capability.</p>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, YOLOx and RT-DETR variants, despite deeper backbones, offer no significant accuracy gains but incur higher GPU memory usage and slower inference, making them less suitable for embedded deployment. Notably, YOLOv11m slightly outperforms YOLOv12m in mAP@0.5 while achieving much higher FPS, confirming its runtime efficiency. More importantly, our proposed RAO-YOLO surpasses all competitors in both mAP@0.5 (<xref ref-type="fig" rid="fig-8">Fig. 8</xref>) and mAP@0.5:0.95, demonstrating superior detection accuracy and precise localization across varying IoU thresholds. This highlights its robustness in challenging RAO scenarios, where fast and reliable hazard identification is critical. Overall, RAO-YOLO achieves the best balance of accuracy, speed, and efficiency, making it the optimal choice for real-world RAO detection tasks.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparative experimental results. The best results in each column are highlighted in bold.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Precision</th>
<th>Recall</th>
<th>mAP@50</th>
<th>mAP@50:95</th>
<th>FPS</th>
<th>Size (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Faster-RCNN [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>0.773</td>
<td>0.591</td>
<td>0.702</td>
<td>0.446</td>
<td>15</td>
<td>42</td>
</tr>
<tr>
<td>RetinaNet [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>0.750</td>
<td>0.680</td>
<td>0.692</td>
<td>0.349</td>
<td>22</td>
<td>38</td>
</tr>
<tr>
<td>RT-DETR-l [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>0.901</td>
<td>0.656</td>
<td>0.734</td>
<td>0.454</td>
<td>111.75</td>
<td>63.1</td>
</tr>
<tr>
<td>RT-DETR-x</td>
<td>0.865</td>
<td>0.665</td>
<td>0.747</td>
<td>0.46</td>
<td>71.19</td>
<td>129.1</td>
</tr>
<tr>
<td>YOLOv9m [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>0.807</td>
<td>0.632</td>
<td>0.756</td>
<td>0.526</td>
<td>240.27</td>
<td>39.0</td>
</tr>
<tr>
<td>YOLOv10m [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>0.869</td>
<td>0.668</td>
<td>0.778</td>
<td>0.533</td>
<td>286.8</td>
<td>31.9</td>
</tr>
<tr>
<td>Hyper-YOLOm [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td><bold>0.903</bold></td>
<td>0.685</td>
<td>0.785</td>
<td>0.532</td>
<td>182.11</td>
<td>59.1</td>
</tr>
<tr>
<td>YOLOv12m [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>0.896</td>
<td>0.688</td>
<td>0.791</td>
<td>0.530</td>
<td>218.51</td>
<td>38.9</td>
</tr>
<tr>
<td>YOLOv12x</td>
<td>0.865</td>
<td><bold>0.693</bold></td>
<td>0.784</td>
<td>0.531</td>
<td>97.89</td>
<td>113.6</td>
</tr>
<tr>
<td>YOLOv11m [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>0.866</td>
<td>0.667</td>
<td>0.792</td>
<td>0.535</td>
<td>261.81</td>
<td>38.6</td>
</tr>
<tr>
<td>YOLOv11x</td>
<td>0.872</td>
<td>0.692</td>
<td>0.790</td>
<td>0.531</td>
<td>122.96</td>
<td>109.1</td>
</tr>
<tr>
<td>RAO-YOLO</td>
<td>0.880</td>
<td>0.643</td>
<td><bold>0.808</bold></td>
<td><bold>0.561</bold></td>
<td>204.04</td>
<td>61.4</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>mAP curves of the ablation experiments.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Ablation Studies</title>
<p>To evaluate the impact of different loss functions on detection performance, we have conducted an ablation study comparing CIoU, Shape-IoU, Inner-IoU, and Focal-MPDIoU within the RAO-YOLO framework. As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, CIoU achieves the highest mAP@0.5 score (0.821), indicating strong performance at loose localization thresholds. However, its mAP@0.5:0.95 score is relatively lower (0.556), suggesting less precise bounding box regression across stricter IoU levels. In contrast, Focal-MPDIoU achieves the highest mAP@0.5:0.95 score (0.561), demonstrating superior localization accuracy under multi-threshold evaluation, which is crucial for detecting small and irregularly shaped RAOs.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of detection performance under different loss functions. The best results in each column are highlighted in bold.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Loss Function</th>
<th>mAP@50</th>
<th>mAP@50:95</th>
</tr>
</thead>
<tbody>
<tr>
<td>CIoU</td>
<td><bold>0.821</bold></td>
<td>0.556</td>
</tr>
<tr>
<td>Shape-IoU</td>
<td>0.800</td>
<td>0.549</td>
</tr>
<tr>
<td>Inner-IoU</td>
<td>0.819</td>
<td>0.551</td>
</tr>
<tr>
<td>Focal-MPDIoU</td>
<td>0.808</td>
<td><bold>0.561</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>While Shape-IoU and Inner-IoU show reasonable performance, their lower mAP metrics indicate less robustness in diverse real-world conditions. The Focal-MPDIoU loss, by combining focal weighting with a geometry-adaptive penalty, effectively mitigates the impact of class imbalance and enhances bounding box refinement. Based on this performance gain, particularly in fine-grained localization, we have adopt Focal-MPDIoU as the default loss function in the final RAO-YOLO model. The comparison of detection performance under different loss functions is presented in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>

<p>To assess the contribution of each component in RAO-YOLO, we conduct ablation studies on the RAOD dataset, focusing on three key modules: MANet, LSDD, and Focal-MPDIoU loss (see <xref ref-type="table" rid="table-4">Table 4</xref>). The baseline YOLOv11m achieves 86.6% precision, 66.7% recall, 79.2% mAP@0.5, and 53.5% mAP@0.5:0.95 at 261.81 FPS (38.6 MB). Adding MANet alone boosts performance to 89.1% precision, 68.5% recall, and 81.4% mAP@0.5, while maintaining &#x003E;200 FPS. Further integrating LSDD raises mAP@0.5 to 82.1% and mAP@0.5:0.95 to 55.6%, albeit with a larger model (61.4 MB) and slightly lower FPS (204.30), enhancing multi-scale RAO detection.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Results of the ablation experiments. The best results in each column are highlighted in bold.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Baseline</th>
<th>MANet</th>
<th>LSDD</th>
<th>Focal-MPDIoU</th>
<th>mAP50</th>
<th>mAP50:95</th>
<th>FPS</th>
<th>Size</th>
</tr>
</thead>
<tbody>
<tr>
<td>&#x2714;</td>
<td></td>
<td></td>
<td></td>
<td>0.792</td>
<td>0.535</td>
<td>261.81</td>
<td>38.6</td>
</tr>
<tr>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td></td>
<td></td>
<td>0.814</td>
<td>0.552</td>
<td>215.16</td>
<td>55.2</td>
</tr>
<tr>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td></td>
<td><bold>0.821</bold></td>
<td>0.556</td>
<td>204.30</td>
<td>61.4</td>
</tr>
<tr>
<td>&#x2714; (Ours)</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>&#x2714;</td>
<td>0.808</td>
<td><bold>0.561</bold></td>
<td>204.04</td>
<td>61.4</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Starting from the YOLOv11m baseline (79.2% mAP@0.5, 53.5% mAP@0.5:0.95), adding MANet improves performance by &#x002B;2.2% mAP@0.5 and &#x002B;1.7% mAP@0.5:0.95; further integrating the LSDD head yields an additional &#x002B;0.7% mAP@0.5 and &#x002B;0.4% mAP@0.5:0.95; finally, the Focal-MPDIoU loss slightlyincreases mAP@0.5:0.95 by &#x002B;0.5%, demonstrating its focus on enhancing high-quality localization. As shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, replacing the original loss with Focal-MPDIoU yields consistently lower training and validation losses, indicating faster convergence and stable performance, confirming its effectiveness in improving RAO detection accuracy.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Training and validation loss curves of the ablation experiments.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-9.tif"/>
</fig>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, RAO-YOLO, which includes all three modules (i.e., MANet, LSDD, and the Focal-MPDIoU loss) achieves the best performance in terms of mAP@0.5:0.95 (56.1%) and maintains strong results in all other metrics. This confirms the effectiveness of the combined enhancements in improving bounding box localization and detection reliability, especially for challenging RAO categories under complex background conditions. These results show that RAO-YOLO&#x2019;s architectural and loss function enhancements work synergistically to achieve high accuracy with acceptable computational efficiency, making it well-suited for real-time roadside perception. This is further illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8</xref> (mAP curves) and <xref ref-type="fig" rid="fig-10">Fig. 10</xref> (PR curves), where our full model consistently outperforms baselines, confirming its superior localization precision.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Precision-Recall (PR) curves of the ablation experiments.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-10.tif"/>
</fig>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Visualization Analysis</title>
<p>To comprehensively evaluate RAO-YOLO, we present heatmaps in <xref ref-type="fig" rid="fig-11">Fig. 11</xref> and qualitative results on representative images in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>, covering diverse real-world conditions including day/night scenes, high-speed traffic, varied weather [<xref ref-type="bibr" rid="ref-50">50</xref>] (e.g., glare, shadows, strong sunlight), and complex road backgrounds (e.g., markings, barriers, curved lanes). Each image includes ground-truth annotations of abandoned objects, which our model accurately localizes with red dotted bounding boxes.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Example heatmap detection results on the RAOD dataset. (<bold>a</bold>) Detection results generated by the YOLOv11m model; (<bold>b</bold>) Detection results generated by the proposed RAO-YOLO model. The heatmaps indicate the response intensity of the models, where warmer colors represent stronger activations. Red bounding boxes highlight the detected RAO targets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-11.tif"/>
</fig><fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Visual comparison of different models on the RAOD dataset. (<bold>a</bold>) Ground truth, where red dashed bounding boxes indicate the locations of RAO targets to be detected. (<bold>b</bold>) Detection results of YOLOv11m, and (<bold>c</bold>) detection results of the proposed RAO-YOLO. In (<bold>b</bold>,<bold>c</bold>), blue bounding boxes denote the detected RAOs. Additionally, in (<bold>b</bold>), red solid bounding boxes highlight missed or inaccurate detections by YOLOv11m. The top-right corner of each image shows a magnified view of the detected region.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-12a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_79851-fig-12b.tif"/>
</fig>
<p>RAO-YOLO shows strong robustness in challenging cases: it reliably detects small, partially occluded, or motion-blurred objects even under low-light conditions. Zoomed-in insets (upper-left corners) highlight its fine-grained localization for small-scale targets which are critical for real-time safety in autonomous driving and roadside monitoring. This visual performance aligns with quantitative gains from our ablation and benchmark studies, driven by the MANet feature extractor, LSDD, and Focal-MPDIoU loss, which together improve boundary precision and detection robustness.</p>
<p>Compared to YOLOv11m, RAO-YOLO produces stronger, more focused activations around true RAOs while suppressing irrelevant background, reflecting enhanced spatial awareness and semantic discrimination. It excels at detecting irregular, elongated, scattered, or occluded RAOs. These improvements stem largely from the Focal-MPDIoU loss, which combines focal weighting with distance-aware IoU to handle class imbalance and geometric variability. This yields tighter bounding boxes, better scale invariance, lower localization error, and greater robustness which are key attributes for deployment in intelligent transportation and edge computing systems where reliability and efficiency are paramount.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper presents RAO-YOLO, a task-specific object detection framework for addressing the critical safety problem of abandoned objects on highways. By integrating the MANet module for enhanced multi-scale feature representation, the LSDD head for efficient and fine-grained localization, and the Focal-MPDIoU loss for robust training on irregular and small targets, RAO-YOLO effectively overcomes the limitations of conventional monocular vision-based methods.</p>
<p>Experiments on the RAOD dataset demonstrate that RAO-YOLO achieves a mAP@0.5:0.95 of 56.1%, outperforming all evaluated baselines while maintaining high inference speed and a compact model size suitable for edge-oriented deployment. Although real-time performance is validated on a GPU platform, future work will further assess its effectiveness on embedded and edge devices. Qualitative results and heatmap visualizations indicate that RAO-YOLO produces more accurate and concentrated responses on abandoned objects, while effectively suppressing background interference such as shadows, road markings, and surrounding vehicles. The model also exhibits robust performance under challenging conditions, including low-light environments, motion blur, occlusion, and irregular object shapes, where baseline methods tend to degrade.</p>
<p>Despite these promising results, several limitations remain. The current model is trained solely on RGB images, which may limit robustness under extreme weather and nighttime conditions. Temporal information from consecutive frames is not explicitly exploited, and evaluation is restricted to a single benchmark. In real-world highway monitoring scenarios, abandoned objects are typically observed in continuous video streams, where temporal cues across frames could provide additional contextual information to improve detection stability and reduce prediction fluctuations. Future work will explore hierarchical RAO categorization (e.g., plastic bags, metal fragments, construction materials), enabling more informative hazard classification for road maintenance operations, and explore synthetic data augmentation to generate additional RAO samples. Future work will address these limitations by incorporating multimodal sensor inputs (e.g., LiDAR, radar), integrating temporal reasoning for video-based detection, and extending validation to more diverse datasets and environments. Such efforts will further enhance the robustness, adaptability, and deployment potential of RAO-YOLO in real-world intelligent transportation systems.</p>
</sec>
</body>
<back>
<ack>
<p>This work is partially supported by the Natural Science Foundation of, China Postdoctoral Science Foundation and Natural Science Foundation of Shandong Province. All the supports are highly appreciated.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work is partially supported by the Natural Science Foundation of China (Grant Number: 52308457), China Postdoctoral Science Foundation (Grant Number: 2024M761811) and Natural Science Foundation of Shandong Province (Grant Number: ZR2023QE220). All the supports are highly appreciated.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization, methodology, writing&#x2014;original draft preparation, Ying Tang; Resources, investigation, writing&#x2014;original draft preparation, Chuanyi Ma; Validation, supervision, writing&#x2014;review and editing, Feng Guo; Software, visualization, Wenhao Sun. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used in this study are publicly available. The RAOD (Road Abandoned Object Detection) dataset can be accessed at: <ext-link ext-link-type="uri" xlink:href="https://github.com/yajunbaby/A-Benchmark-for-Road-Abandoned-Object-Detection-from-Video-Surveillance">https://github.com/yajunbaby/A-Benchmark-for-Road-Abandoned-Object-Detection-from-Video-Surveillance</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>NHTSA</collab></person-group>. <article-title>Administration, motor vehicle traffic crash data resource page, national highway traffic safety administration</article-title>. <year>2022 [cited 2026 Jan 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://crashstats.nhtsa.dot.gov/#!/">https://crashstats.nhtsa.dot.gov/#!/</ext-link>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Cost analysis of road traffic crashes in China</article-title>. <source>Int J Inj Control Saf Promot</source>. <year>2020</year>;<volume>27</volume>(<issue>3</issue>):<fpage>385</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1080/17457300.2020.1785507</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>FR</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ning</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Big data analytics in intelligent transportation systems: a survey</article-title>. <source>IEEE Trans Intell Transport Syst</source>. <year>2019</year>;<volume>20</volume>(<issue>1</issue>):<fpage>383</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2018.2815678</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bebis</surname> <given-names>G</given-names></string-name>, <string-name><surname>Miller</surname> <given-names>R</given-names></string-name></person-group>. <article-title>On-road vehicle detection: a review</article-title>. <source>IEEE Trans Pattern Anal Machine Intell</source>. <year>2006</year>;<volume>28</volume>(<issue>5</issue>):<fpage>694</fpage>&#x2013;<lpage>711</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2006.104</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hoffmann</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Tosso</surname> <given-names>HG</given-names></string-name>, <string-name><surname>Santos</surname> <given-names>MMD</given-names></string-name>, <string-name><surname>Justo</surname> <given-names>JF</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>AW</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>AU</given-names></string-name></person-group>. <article-title>Real-time adaptive object detection and tracking for autonomous vehicles</article-title>. <source>IEEE Trans Intell Veh</source>. <year>2021</year>;<volume>6</volume>(<issue>3</issue>):<fpage>450</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tiv.2020.3037928</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Stauffer</surname> <given-names>C</given-names></string-name>, <string-name><surname>Grimson</surname> <given-names>WE</given-names></string-name></person-group>. <article-title>Adaptive background mixture models for real-time tracking</article-title>. In: <conf-name>Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR); 1999 Jun 23&#x2013;25</conf-name>; <publisher-loc>Fort Collins, CO, USA</publisher-loc>. p. <fpage>246</fpage>&#x2013;<lpage>52</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kass</surname> <given-names>M</given-names></string-name>, <string-name><surname>Witkin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Terzopoulos</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Snakes: active contour models</article-title>. <source>Int J Comput Vis</source>. <year>1988</year>;<volume>1</volume>(<issue>4</issue>):<fpage>321</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1007/BF00133570</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Champahom</surname> <given-names>T</given-names></string-name>, <string-name><surname>Se</surname> <given-names>C</given-names></string-name>, <string-name><surname>Watcharamaisakul</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jomnonkwao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Karoonsoontawong</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ratanavaraha</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Tree-based approaches to understanding factors influencing crash severity across roadway classes: a Thailand case study</article-title>. <source>IATSS Res</source>. <year>2024</year>;<volume>48</volume>(<issue>3</issue>):<fpage>464</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.iatssr.2024.09.001</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Labayrade</surname> <given-names>R</given-names></string-name>, <string-name><surname>Aubert</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tarel</surname> <given-names>JP</given-names></string-name></person-group>. <article-title>Real time obstacle detection in stereovision on non flat road geometry through "v-disparity" representation</article-title>. In: <conf-name>Proceedings of the Intelligent Vehicle Symposium; 2002 Jun 17&#x2013;21</conf-name>; <publisher-loc>Versailles, France</publisher-loc>. p. <fpage>646</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matthies</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shafer</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Error modeling in stereo navigation</article-title>. <source>IEEE J Robot Automat</source>. <year>1987</year>;<volume>3</volume>(<issue>3</issue>):<fpage>239</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jra.1987.1087097</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Multi-view 3D object detection network for autonomous driving</article-title>. In: <conf-name>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>1907</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2017.691</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A survey of 3D object detection algorithms for intelligent vehicles development</article-title>. <source>Artif Life Robot</source>. <year>2022</year>;<volume>27</volume>(<issue>1</issue>):<fpage>115</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10015-021-00711-0</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Faster R-CNN towards real-time object detection with region proposal networks</article-title>. In: <conf-name>Advances in neural information processing systems</conf-name>. <publisher-loc>Montr&#x00E9;al, QC, Canada</publisher-loc>; <year>2015</year>. p. <fpage>91</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Anguelov</surname> <given-names>D</given-names></string-name>, <string-name><surname>Erhan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Reed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>CY</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>SSD: single shot multibox detector</chapter-title>. In: <source>Computer vision&#x2014;ECCV 2016</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2016</year>. p. <fpage>21</fpage>&#x2013;<lpage>37</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Redmon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Divvala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Farhadi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>You only look once: unified, real-time object detection</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2016.91</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Veres</surname> <given-names>M</given-names></string-name>, <string-name><surname>Moussa</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deep learning for intelligent transportation systems: a survey of emerging trends</article-title>. <source>IEEE Trans Intell Transport Syst</source>. <year>2020</year>;<volume>21</volume>(<issue>8</issue>):<fpage>3152</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2019.2929020</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cordts</surname> <given-names>M</given-names></string-name>, <string-name><surname>Omran</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramos</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rehfeld</surname> <given-names>T</given-names></string-name>, <string-name><surname>Enzweiler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Benenson</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The cityscapes dataset for semantic urban scene understanding</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>3213</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2016.350</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Geiger</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lenz</surname> <given-names>P</given-names></string-name>, <string-name><surname>Urtasun</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Are we ready for autonomous driving? The KITTI vision benchmark suite</article-title>. In: <conf-name>Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition; 2012 Jun 16&#x2013;21</conf-name>; <publisher-loc>Providence, RI, USA</publisher-loc>. p. <fpage>3354</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2012.6248074</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pinggera</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramos</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gehrig</surname> <given-names>S</given-names></string-name>, <string-name><surname>Franke</surname> <given-names>U</given-names></string-name>, <string-name><surname>Rother</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mester</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Lost and found: detecting small road hazards for self-driving vehicles</article-title>. In: <conf-name>Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2016 Oct 9&#x2013;14</conf-name>; <publisher-loc>Daejeon, Republic of Korea</publisher-loc>. p. <fpage>1099</fpage>&#x2013;<lpage>106</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iros.2016.7759186</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Radke</surname> <given-names>RJ</given-names></string-name></person-group>. <article-title>Context-aware video anomaly detection in long-term datasets</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2024; 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>4002</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW63382.2024.00404</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xian</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>BDD100K: a diverse driving dataset for heterogeneous multitask learning</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>2636</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.00271</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lis</surname> <given-names>K</given-names></string-name>, <string-name><surname>Uhlemeyer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Blum</surname> <given-names>H</given-names></string-name>, <string-name><surname>Honari</surname> <given-names>S</given-names></string-name>, <string-name><surname>Siegwart</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Segmentmeifyoucan: a benchmark for anomaly segmentation</article-title>. <comment>arXiv:2104.14812. 2021</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Nan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RAOD: a benchmark for road abandoned object detection from video surveillance</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>123985</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3407955</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Latha</surname> <given-names>YM</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>BS</given-names></string-name></person-group>. <chapter-title>A systematic review on background subtraction model for data detection</chapter-title>. In: <source>Pervasive computing and social networking</source>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>341</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-16-5640-8_27</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ojala</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pietik&#x00E4;inen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Harwood</surname> <given-names>D</given-names></string-name></person-group>. <article-title>A comparative study of texture measures with classification based on featured distributions</article-title>. <source>Pattern Recognit</source>. <year>1996</year>;<volume>29</volume>(<issue>1</issue>):<fpage>51</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0031-3203(95)00067-4</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dalal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Triggs</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Histograms of oriented gradients for human detection</article-title>. In: <conf-name>Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR&#x2019;05); 2005 Jun 20&#x2013;26</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>. p. <fpage>886</fpage>&#x2013;<lpage>93</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garcia-Garcia</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bouwmans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rosales Silva</surname> <given-names>AJ</given-names></string-name></person-group>. <article-title>Background subtraction in real applications: challenges, current models and future directions</article-title>. <source>Comput Sci Rev</source>. <year>2020</year>;<volume>35</volume>(<issue>1</issue>):<fpage>100204</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cosrev.2019.100204</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ros</surname> <given-names>G</given-names></string-name>, <string-name><surname>Sellart</surname> <given-names>L</given-names></string-name>, <string-name><surname>Materzynska</surname> <given-names>J</given-names></string-name>, <string-name><surname>Vazquez</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lopez</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>The SYNTHIA dataset: a large collection of synthetic images for semantic segmentation of urban scenes</article-title>. In: <conf-name>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>3234</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2016.352</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bakirci</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Advanced aerial monitoring and vehicle classification for intelligent transportation systems with YOLOv8 variants</article-title>. <source>J Netw Comput Appl</source>. <year>2025</year>;<volume>237</volume>(<issue>B</issue>):<fpage>104134</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jnca.2025.104134</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Carion</surname> <given-names>N</given-names></string-name>, <string-name><surname>Massa</surname> <given-names>F</given-names></string-name>, <string-name><surname>Synnaeve</surname> <given-names>G</given-names></string-name>, <string-name><surname>Usunier</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kirillov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zagoruyko</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>End-to-end object detection with transformers</chapter-title>. In: <source>Computer vision&#x2014;ECCV 2020</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2020</year>. p. <fpage>213</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-58452-8_13</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Su</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deformable DETR: deformable transformers for end-to-end object detection</article-title>. <comment>arXiv:2010.04159. 2021</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>F</given-names></string-name></person-group>. <article-title>D-FINE: redefine regression task in DETRs as fine-grained distribution refinement</article-title>. <comment>arXiv:2410.13842. 2024</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DETRs beat YOLOs on real-time object detection</article-title>. <comment>arXiv:2304.08069. 2023</comment>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ku</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mozifian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Harakeh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Waslander</surname> <given-names>SL</given-names></string-name></person-group>. <article-title>Joint 3D proposal generation and object detection from view aggregation</article-title>. In: <conf-name>Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2018 Oct 1&#x2013;5</conf-name>; <publisher-loc>Madrid, Spain</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iros.2018.8594049</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Azadvatan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kurt</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Melnet: a real-time deep learning algorithm for object detection</article-title>. <comment>arXiv:2401.17972. 2024</comment>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>TinyDet: accurate small object detection in lightweight generic detectors</article-title>. <comment>arXiv:2304.03428. 2023</comment>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Woo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Kweon</surname> <given-names>IS</given-names></string-name></person-group>. <chapter-title>CBAM: convolutional block attention module</chapter-title>. In: <source>Computer vision&#x2014;ECCV 2018</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2018</year>. p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_1</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>W</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>ECA-net: efficient channel attention for deep convolutional neural networks</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>11531</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.01155</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Du</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ying</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yong</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hyper-YOLO: when visual object detection meets hypergraph computation</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2025</year>;<volume>47</volume>(<issue>4</issue>):<fpage>2388</fpage>&#x2013;<lpage>401</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2024.3524377</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name></person-group>. <chapter-title>Group normalization</chapter-title>. In: <source>Computer vision&#x2014;ECCV 2018</source>. <publisher-loc>Cham, Swizerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2018</year>. p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01261-8_1</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Distance-IoU loss: faster and better learning for bounding box regression</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>(<issue>7</issue>):<fpage>12993</fpage>&#x2013;<lpage>3000</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i07.6999</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Focaler-IoU: more focused intersection over union loss</article-title>. <comment>arXiv:2401.10525. 2024</comment>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>MPDIoU: a loss for efficient and accurate bounding box regression</article-title>. <comment>arXiv:2307.07662. 2023</comment>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Padilla</surname> <given-names>R</given-names></string-name>, <string-name><surname>Netto</surname> <given-names>SL</given-names></string-name>, <string-name><surname>da Silva</surname> <given-names>EAB</given-names></string-name></person-group>. <article-title>A survey on performance metrics for object-detection algorithms</article-title>. In: <conf-name>Proceedings of the 2020 International Conference on Systems, Signals and Image Processing (IWSSIP); 2020 Jul 1&#x2013;3</conf-name>; <publisher-loc>Niter&#x00F3;i, Brazil</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/iwssip48289.2020.9145130</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Dollar</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Focal loss for dense object detection</article-title>. In: <conf-name>Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV); 2017 Oct 22&#x2013;29</conf-name>; <publisher-loc>Venice, Italy</publisher-loc>. p. <fpage>2980</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2017.324</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Yeh</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <article-title>YOLOv9: learning what you want to learn using programmable gradient information</article-title>. <comment>arXiv:2402.13616. 2024</comment>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>YOLOv10: real-time end-to-end object detection</article-title>. <comment>arXiv:2405.14458. 2024</comment>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Doermann</surname> <given-names>D</given-names></string-name></person-group>. <article-title>YOLOv12: attention-centric real-time object detectors</article-title>. <comment>arXiv:2502.12524. 2025</comment>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Khanam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>M</given-names></string-name></person-group>. <article-title>YOLOv11: an overview of the key architectural enhancements</article-title>. <comment>arXiv:2410.17725. 2024</comment>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iqra</surname></string-name>, <string-name><surname>Giri</surname> <given-names>KJ</given-names></string-name>, <string-name><surname>Javed</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Small object detection in diverse application landscapes: a survey</article-title>. <source>Multimed Tools Appl</source>. <year>2024</year>;<volume>83</volume>(<issue>41</issue>):<fpage>88645</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-024-18866-w</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>





