<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">JAI</journal-id>
<journal-id journal-id-type="nlm-ta">JAI</journal-id>
<journal-id journal-id-type="publisher-id">JAI</journal-id>
<journal-title-group>
<journal-title>Journal on Artificial Intelligence</journal-title>
</journal-title-group>
<issn pub-type="epub">2579-003X</issn>
<issn pub-type="ppub">2579-0021</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">76274</article-id>
<article-id pub-id-type="doi">10.32604/jai.2026.076274</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>EMA-GhostConv YOLOv8 Based Enhanced Vehicle Detection in Intelligent Transportation Applications</article-title>
<alt-title alt-title-type="left-running-head">EMA-GhostConv YOLOv8 Based Enhanced Vehicle Detection in Intelligent Transportation Applications</alt-title>
<alt-title alt-title-type="right-running-head">EMA-GhostConv YOLOv8 Based Enhanced Vehicle Detection in Intelligent Transportation Applications</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Rahman</surname><given-names>A. S. M. Masudur</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Zamir</surname><given-names>Muhammad Zunair</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ullah</surname><given-names>Syed Sajid</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>sajid@chd.edu.cn</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Khan</surname><given-names>Salman</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Saman</surname><given-names>Maria</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Bahadar</surname><given-names>Naqash</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information Engineering, Chang&#x2019;an University</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Energy and Electrical Engineering, Chang&#x2019;an University</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Syed Sajid Ullah. Email: <email>sajid@chd.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>24</day><month>2</month><year>2026</year>
</pub-date>
<volume>8</volume>
<issue>1</issue>
<fpage>119</fpage>
<lpage>136</lpage>
<history>
<date date-type="received">
<day>17</day>
<month>11</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_JAI_76274.pdf"></self-uri>
<abstract>
<p>Vehicle detection plays a pivotal role in autonomous driving, traffic monitoring, and intelligent surveillance systems. While YOLOv8 offers strong real-time performance, its detection accuracy is often limited by insufficient feature stability and suboptimal multi-scale feature fusion in complex scenes. To address these issues, we propose an enhanced YOLOv8 framework that retains the original backbone and detection head for efficiency while introducing targeted improvements to the neck architecture. Specifically, the model incorporates an Exponential Moving Average (EMA) feature layer to stabilize learning through temporally smoothed feature representations, which reduces noise and enhances generalization, and integrates GhostConv along with C3Ghost modules to enrich feature diversity with minimal computational overhead. We conduct comprehensive experiments on both a custom vehicle dataset and the KITTI benchmark, demonstrating consistent and significant improvements over the baseline YOLOv8. On KITTI, our method achieves gains of &#x002B;11.68% in precision, &#x002B;12.84% in recall, &#x002B;10.20% in mAP@50, and &#x002B;12.73% in mAP@50&#x2013;95. Furthermore, ablation studies validate the individual contribution of each component, and comparisons with multiple state-of-the-art detectors confirm the competitiveness of our approach in both accuracy and efficiency. The proposed EMA-GhostConv YOLOv8 framework thus offers a robust, lightweight, and high-performance solution for vehicle detection in intelligent transportation applications.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>YOLOv8n</kwd>
<kwd>vehicle detection</kwd>
<kwd>deformable convolutional networks (DCNv2)</kwd>
<kwd>ghost module</kwd>
<kwd>exponential moving average (EMA)</kwd>
<kwd>attention mechanisms</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The rapid evolution of intelligent transportation systems (ITS) has necessitated the development of highly accurate and efficient real-time object detection models for tasks such as vehicle identification, traffic monitoring, and autonomous driving. These applications demand lightweight architectures that can operate under strict latency and resource constraints, especially in embedded or edge environments. Object detection, as a foundational task in computer vision, has witnessed significant breakthroughs driven by deep convolutional neural networks (CNNs), among which the <italic>You Only Look Once (YOLO)</italic> family of models stands out for balancing speed and accuracy.</p>
<p>The YOLO paradigm emerged with YOLOv1 [<xref ref-type="bibr" rid="ref-1">1</xref>], introducing a unified, single-stage architecture capable of real-time detection. This was further improved in YOLOv2 [<xref ref-type="bibr" rid="ref-2">2</xref>] and YOLOv3 [<xref ref-type="bibr" rid="ref-3">3</xref>] through enhancements in backbone networks and anchor-based predictions. YOLOv4 [<xref ref-type="bibr" rid="ref-4">4</xref>] leveraged additional tricks such as CSPDarknet53 and bag-of-freebies techniques to increase mAP without sacrificing speed. While YOLOv5 [<xref ref-type="bibr" rid="ref-5">5</xref>] did not emerge from a peer-reviewed outlet, it became widely adopted for its modularity and deployment flexibility. YOLOv6 [<xref ref-type="bibr" rid="ref-6">6</xref>] focused on industrial scenarios with anchor-free mechanisms, and YOLOv7 [<xref ref-type="bibr" rid="ref-7">7</xref>] introduced trainable bag-of-freebies and E-ELAN modules for efficient training.</p>
<p>The most recent versions, YOLOv8 [<xref ref-type="bibr" rid="ref-8">8</xref>], YOLOv9 [<xref ref-type="bibr" rid="ref-9">9</xref>], and YOLOv10 [<xref ref-type="bibr" rid="ref-10">10</xref>], integrate novel strategies such as native Python implementation, programmable gradient information, and end-to-end detection pipelines. These developments reflect the community&#x2019;s shift toward optimizing both model architecture and training efficiency. Despite these advances, challenges remain in improving detection in cluttered, occluded, and dynamically varying environments typical in real-world ITS. Additionally, computational burdens imposed by dense feature extraction and large-scale convolutional layers often hinder deployment on resource-constrained platforms.</p>
<p>To address these limitations, recent research has explored advanced architectural modules aimed at reducing computational complexity while maintaining detection precision. Yadav et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] incorporated Ghost convolution and GhostBottleneck layers into YOLOv5s to reduce redundancy in feature extraction. Similarly, Bao and Gao [<xref ref-type="bibr" rid="ref-12">12</xref>] introduced YED-YOLO for autonomous driving, optimizing the detection pipeline using lightweight convolutional mechanisms. Du et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] proposed GhostC2f-enhanced YOLOv8 models for distracted driving detection, demonstrating real-time performance in resource-constrained environments. Lv et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] further enhanced model efficiency by integrating GhostNetV2 with SE attention in GS-YOLO, while Tang et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] extended the GhostNet architecture to support long-range attention with minimal overhead.</p>
<p>In addition to architectural modifications, optimization techniques have also shown promise. Liu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] designed a lightweight vehicle detection model suitable for real-time deployment in ITS. Ullah et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] augmented YOLOv8 with GhostConv and attention modules for improved performance in dynamic traffic scenarios. The use of Exponential Moving Average (EMA) was explored by Li et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] in EMA-YOLO to stabilize training and boost convergence, and similarly by Han et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] in a maritime object detection context. Zhang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] leveraged multi-coordinate aggregation attention and shared convolution for fast and accurate road defect detection. Collectively, these studies demonstrate that incorporating Ghost-based convolutions, attention mechanisms, and EMA techniques can significantly improve detection performance, training stability, and computational efficiency in intelligent transportation systems.</p>
<p>In this context, we propose a novel EMA-GhostConv-enhanced YOLOv8 framework tailored for robust vehicle detection in ITS environments. By embedding GhostConv layers within the YOLOv8 backbone and integrating EMA-based parameter updates, our model achieves a better trade-off between speed and accuracy. Experimental evaluations on benchmark traffic datasets validate the efficacy of the proposed method, outperforming existing state-of-the-art lightweight detectors in both precision and inference time.</p>
<p>The key contributions of this paper are as follows:
<list list-type="order">
<list-item>
<p>To enhance the neck of YOLOv8n, we introduce a novel EMA feature layer for stabilizing multi-scale feature learning alongside a lightweight architecture built by replacing standard convolutions with unified GhostConv-C3Ghost blocks. This dual approach improves generalization against noise and significantly reduces computational redundancy while preserving expressive power for accurate detection in complex scenes.</p></list-item>
<list-item>
<p>We conduct a rigorous ablation study to quantify the individual and combined impact of the EMA and GhostConv components, demonstrating their complementary roles in boosting precision, recall, and mAP.</p></list-item>
<list-item>
<p>We conduct extensive evaluations on both the standard KITTI benchmark and a challenging custom vehicle dataset collected under diverse real-world conditions&#x2014;such as partial occlusions, varying illumination, and complex urban backgrounds. The proposed EMA-GhostConv YOLOv8 model is comprehensively compared with multiple state-of-the-art detectors, including various YOLO variants and Faster R-CNN, SSD, RetinaNet, DETR, and EfficientDet-D2. Experimental results demonstrate that our method consistently outperforms all baselines in terms of detection accuracy and robustness.</p></list-item>
</list></p>
<p>The overall structure of the paper is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The remainder of this paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews related work in vehicle detection and deep learning architectures. <xref ref-type="sec" rid="s3">Section 3</xref> presents the proposed methodology. <xref ref-type="sec" rid="s4">Section 4</xref> details the experimental setup and results. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper and outlines potential future research directions.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Structural outline of the paper, illustrating the organization of sections from introduction to conclusion.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Real-time object detection in intelligent transportation systems (ITS) presents persistent challenges, particularly in balancing computational efficiency with detection accuracy under complex conditions. A wide range of YOLO-based architectures has been adapted for ITS scenarios, with recent efforts shifting toward task-specific lightweight designs and embedded deployment. Liu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed a compact detection model for real-time vehicle recognition on resource-limited hardware. Bakirci [<xref ref-type="bibr" rid="ref-21">21</xref>] explored aerial monitoring using YOLOv8 variants to classify vehicles across road networks. Zeng and Zhong [<xref ref-type="bibr" rid="ref-22">22</xref>] targeted road surface anomalies using a customized YOLOv8n variant, while Zhu et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] enhanced infrastructure inspection through UAV-mounted YOLOv8 detectors. Du et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] integrated GhostC2f modules within YOLOv8 to detect distracted driving with high speed and accuracy.</p>
<p>To reduce computational cost without sacrificing performance, Ghost convolution and its variants have become popular. Tang et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] introduced GhostNetv2 to minimize redundant feature maps while incorporating long-range attention. Lv et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] built on this by combining GhostNetV2 with SE attention in GS-YOLO, achieving superior performance in SAR ship detection. Yadav et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] embedded GhostConv and GhostBottleneck layers in YOLOv5s to reduce latency while maintaining high precision. Ullah et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] applied GhostConv and attention into YOLOv8 for dense ITS vehicle detection, while Elhenidy et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] introduced ghost separable convolution in GY-YOLO for pedestrian detection.</p>
<p>The use of Exponential Moving Average (EMA) has shown potential in stabilizing training and improving model generalization. Li et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed EMA-YOLO, enhancing object detection consistency across dynamic scenes. Han et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] also incorporated EMA with channel attention in Light-YOLOv7 to detect maritime targets efficiently. These contributions demonstrate that EMA is a powerful yet underutilized technique in lightweight detection.</p>
<p>Several studies have also fused attention modules to strengthen detection in low-contrast and occluded scenarios. Zhang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] developed a multi-coordinate aggregation attention block for road defect detection using shared convolution, and Wei et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] introduced a multi-scale attention-enhanced YOLO model for remote sensing. Zhou et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] combined layer-adaptive pruning and multidimensional feature fusion in Mp-YOLO for dense traffic object detection.</p>
<p>In addition, a range of YOLO-based models has been customized for ITS-specific applications. Thatikonda [<xref ref-type="bibr" rid="ref-27">27</xref>] built a lightweight helmet and license plate detector using YOLOv8. Bao and Gao [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed YED-YOLO for autonomous driving, demonstrating a lightweight yet accurate solution. Yu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] optimized YOLOv8 for general vehicle detection across environments, and Pan et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] introduced LVD-YOLO for low-resource ITS deployments.</p>
<p>Despite these valuable efforts, most existing models address only partial aspects of the challenges, such as computational overhead, unstable training, or limited generalization, without integrating all three effectively. As summarized in <xref ref-type="table" rid="table-1">Table 1</xref>, prior approaches typically optimize for one or two factors, but few offer a unified solution. Our proposed EMA-GhostConv YOLOv8 model is designed to overcome these gaps by unifying Ghost convolution, EMA training stabilization, and lightweight attention into a single, coherent architecture. The result is a highly accurate, computationally efficient, and robust detector tailored for real-time ITS vehicle detection in complex environments.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparative summary of YOLO-based detection models.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Ref.</th>
<th>Year</th>
<th>Model/Method</th>
<th>Limitations</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>2023</td>
<td>YOLOv5s &#x002B; GhostConv</td>
<td>Uses older YOLOv5; limited adaptability to complex scenes.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>2024</td>
<td>YED-YOLO</td>
<td>Domain-specific; not validated on diverse datasets.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>2023</td>
<td>YOLOv8 &#x002B; GhostC2f</td>
<td>Focused on efficiency; limited contextual fusion.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>2024</td>
<td>GS-YOLO (GhostNetV2 &#x002B; SE)</td>
<td>SAR-specific; weak cross-domain generalization.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>2024</td>
<td>Lightweight YOLO</td>
<td>Prioritizes compactness over detection robustness.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>2025</td>
<td>YOLOv8 &#x002B; GhostConv &#x002B; Attention</td>
<td>Inconsistent multi-scale feature learning.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>2025</td>
<td>EMA-YOLO</td>
<td>Stable but lacks lightweight adaptation.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>2024</td>
<td>Light-YOLOv7 &#x002B; EMA</td>
<td>Efficient yet weak on small-object detection.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>2025</td>
<td>YOLO &#x002B; Attention &#x002B; Shared Conv</td>
<td>Task-specific; limited generalization.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>2025</td>
<td>YOLOv8 Variants</td>
<td>Aerial-focused; no real-world validation.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>2024</td>
<td>YOLOv8n Improvement</td>
<td>Optimized for road damage only.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>2024</td>
<td>YOLOv8 &#x002B; UAV</td>
<td>UAV-oriented; moderate precision in complex lighting.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>2025</td>
<td>GY-YOLO</td>
<td>Pedestrian-specific; poor multi-class scalability.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>2025</td>
<td>SED-YOLO</td>
<td>Remote-sensing focus; low inference efficiency.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>2024</td>
<td>Mp-YOLO &#x002B; Pruning</td>
<td>Compact but reduced accuracy on small targets.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>2024</td>
<td>YOLOv8 Lightweight</td>
<td>Helmet/license focus; limited general use.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>2024</td>
<td>Improved YOLOv8</td>
<td>Minor tuning; lacks architectural novelty.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>2024</td>
<td>LVD-YOLO</td>
<td>Low-resource design; weak in dense urban scenes.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>This section outlines the proposed vehicle detection framework built upon the YOLOv8 object detection model. We first present an overview of the original YOLOv8 architecture, followed by the integration of our key enhancements: Exponential Moving Average (EMA) for stabilized learning, Ghost convolution for lightweight feature extraction, and a C3Ghost module to support multi-scale representation. The modified framework is trained and evaluated on two datasets, KITTI and a custom real-world traffic dataset, designed to test its performance under diverse road conditions. We also describe the metrics used to evaluate the detection performance.</p>
<sec id="s3_1">
<label>3.1</label>
<title>YOLOv8 Architecture</title>
<p>The baseline YOLOv8 framework follows a single-stage, anchor-free object detection paradigm and consists of three core components: a CSP-based backbone that extracts hierarchical features across multiple spatial resolutions; a neck employing a Path Aggregation Network with Feature Pyramid Network (PAN-FPN) to fuse low-level and high-level features for multi-scale representation; and a decoupled detection head that separately handles classification and regression to directly output object categories and bounding box coordinates. While YOLOv8 achieves a strong trade-off between accuracy and speed, its neck architecture suffers from limited gradient stability and inadequate cross-scale feature interaction, particularly when detecting small or partially occluded vehicles (see <xref ref-type="fig" rid="fig-2">Fig. 2</xref>). To overcome these limitations, this work introduces targeted architectural enhancements that improve feature stability and multi-scale representation while preserving real-time inference capability.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Overall architecture of the YOLOv8.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Proposed Modification</title>
<p>To address the limitations identified in recent YOLOv8-based vehicle detection frameworks, we propose a set of architectural modifications aimed at enhancing both detection stability and computational efficiency for intelligent transportation systems. While YOLOv8 provides a strong baseline due to its anchor-free detection and decoupled head architecture, it often struggles in scenarios involving multi-scale objects, illumination variations, and partial occlusion conditions commonly encountered in real-world traffic environments. Furthermore, the demand for deployment on edge or embedded platforms necessitates the development of lightweight yet robust detection mechanisms.</p>
<p>Our proposed architecture enhances YOLOv8 by integrating three core modules: (i) an Exponential Moving Average (EMA) layer for stabilized weight updates and improved generalization; (ii) Ghost convolution (GhostConv) to reduce computational redundancy while preserving representational power; and (iii) a modified C3Ghost module for enriched multi-scale feature extraction with minimal inference overhead. Together, these components create a hybrid model that improves detection accuracy under challenging conditions without compromising real-time performance. The following subsections detail each of these modifications.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Exponential Moving Average (EMA)</title>
<p>To enhance the stability of feature learning and suppress overfitting during training, an Exponential Moving Average (<italic>EMA</italic>) mechanism is integrated into the neck component of the proposed architecture. <italic>EMA</italic> enables temporal smoothing of feature map updates, thereby mitigating the influence of noisy gradients, illumination fluctuations, and abrupt structural changes, conditions often present in real-world traffic scenes.</p>
<p>At each training iteration <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>t</mml:mi></mml:math></inline-formula>, the smoothed feature representation <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is updated as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>here, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>F</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> denotes the current feature map, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the previous smoothed estimate, and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is the decay factor controlling the balance between current and past observations. A smaller <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> places more emphasis on historical data, promoting stability, while a larger <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> allows faster adaptation to recent changes.</p>
<p>By continuously updating features in a smoothed manner, the EMA module enhances the model&#x2019;s robustness against motion blur, partial occlusions, and sensor noise, ultimately contributing to improved detection accuracy under dynamic road conditions.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>GhostConv</title>
<p>Ghost convolution (GhostConv) is a lightweight convolutional module designed to reduce redundancy in feature extraction by generating more feature maps using inexpensive linear operations. Instead of relying solely on standard convolution, GhostConv decomposes the feature generation process into two stages: a small set of intrinsic features are produced via conventional convolution, while the remaining features, termed ghost features, are obtained through computationally efficient linear transformations.</p>
<p>Mathematically, the GhostConv operation can be expressed as:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the primary feature maps generated by standard convolution, and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents a set of linear operations (e.g., depthwise convolutions or pointwise filters) applied to generate the ghost features. The concatenation <italic>Y</italic> forms the final output with significantly reduced computational cost.</p>
<p>By replacing standard convolutions in the neck of YOLOv8 with GhostConv blocks, the model achieves lower FLOPs and parameter count while maintaining sufficient representational power for object detection tasks.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>C3Ghost</title>
<p>The C3Ghost module is an architectural variant of the original C3 block in YOLOv5/YOLOv8, modified to incorporate GhostConv layers as its internal convolutional units. It consists of a residual bottleneck structure, where GhostConv is employed in both the main and skip branches to ensure computational efficiency across the entire module.</p>
<p>C3Ghost retains the multi-branch feature aggregation capability of the standard C3 module, enabling it to effectively combine deep and shallow features. This promotes richer contextual learning and improved multi-scale feature fusion, which are critical in detecting small and partially occluded vehicles in complex traffic environments.</p>
<p>While it does not introduce new mathematical operations beyond GhostConv, C3Ghost amplifies the benefits of ghost-based computation within a residual learning framework. When embedded within the YOLOv8 neck, the C3Ghost module contributes to both improved feature abstraction and faster inference, aligning with real-time constraints of intelligent transportation systems.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Proposed Architecture and Algorithmic Workflow</title>
<p>The proposed detection framework builds upon the original YOLOv8 design. Its neck architecture integrates the Exponential Moving Average (EMA) for enhanced feature stability and a unified GhostConv-C3Ghost block, where GhostConv and C3Ghost are integrated as one model, to achieve reduced computational cost and improved multi-scale detection performance.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the architecture is composed of three primary stages: Backbone, Neck, and Head.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Overall architecture of the proposed EMA-GhostConv YOLOv8 framework. The backbone and detection head are retained from YOLOv8, while the neck is enhanced using EMA and GhostConv modules.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-3.tif"/>
</fig>
<p><list list-type="simple">
<list-item><label>1)</label><p><bold>Backbone:</bold> The backbone retains the core YOLOv8 feature extractor, comprising convolutional blocks (CBS), Cross Stage Partial blocks (C2f), and a spatial pyramid pooling fast (SPPF) module. Feature maps are extracted at five pyramid levels, denoted as <italic>P1</italic> through <italic>P5</italic>, with progressively decreasing spatial resolutions and increasing channel depths. These layers encode both spatial and semantic information crucial for object localization.</p></list-item>
</list></p>
<p><list list-type="simple">
<list-item><label>2)</label><p><bold>Neck:</bold> In the neck, we introduce two upsampling paths to refine multi-scale features from <italic>P5</italic> and <italic>P4</italic>:</p>
<p>- In the bottom-up path, feature maps from deeper layers (e.g., <italic>P5</italic>) are upsampled and concatenated with corresponding shallow features (e.g., <italic>P4</italic>), then passed through C3Ghost modules. - EMA modules are inserted after each major fusion step to stabilize the evolving feature representations during training, thus improving robustness under noisy or dynamic conditions. - GhostConv layers replace standard convolutions in the fusion blocks to reduce parameter count and floating point operations (FLOPs), ensuring that the model remains computationally efficient.</p></list-item>
</list></p>
<p><list list-type="simple">
<list-item><label>3)</label><p><bold>Detection Head:</bold> The fused and smoothed feature maps at different scales (<italic>P3, P4, P5</italic>) are forwarded to independent decoupled detection heads, which perform object classification and bounding box regression. This design supports the detection of vehicles at multiple scales, including small and partially occluded objects.</p></list-item>
</list></p>
<p>Overall, the architectural modifications enable enhanced feature abstraction while maintaining low latency, making the model suitable for real-time deployment in intelligent transportation systems. The use of both KITTI and a custom dataset ensures generalization across benchmark and real-world environments.</p>
<p>Algorithm 1 outlines the main steps of the proposed EMA-GhostConv YOLOv8 training pipeline.</p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-8.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Datasets</title>
<p>To evaluate the performance and generalization of the proposed EMA-GhostConv YOLOv8 framework, we used two complementary datasets: the public KITTI benchmark and a custom real-world vehicle dataset. The KITTI dataset provides high-quality, annotated traffic scenes under diverse but standardized conditions and was used for binary vehicle detection, where the unified &#x201C;vehicle&#x201D; class encompasses cars, buses, trucks, and vans. Images were resized and split 80:20 into training and validation sets.</p>
<p>The custom dataset was collected from urban and semi-urban environments across day, dusk, and night conditions, featuring challenges like occlusions, motion blur, and non-standard vehicle appearances. All images were manually annotated in YOLO format and augmented to enhance generalization. Together, these datasets enable rigorous evaluation under both benchmarked and realistic, unconstrained traffic scenarios.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Evaluation Metrics</title>
<p>The performance of the proposed EMA-GhostConv YOLOv8 model is quantitatively assessed using four standard evaluation metrics widely adopted in object detection: Precision, Recall, and mean Average Precision (mAP) [<xref ref-type="bibr" rid="ref-30">30</xref>]. These metrics jointly reflect both the classification accuracy and the spatial localization capability of the model. Definitions and interpretations of these metrics are summarized in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of key evaluation metrics for object detection.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Definition &#x0026; Interpretation</th>
<th>Formula</th>
</tr>
</thead>
<tbody>
<tr>
<td>Precision (Positive Predictive Value)</td>
<td>Proportion of correctly predicted positives among all predicted positives. High precision indicates fewer false alarms.</td>
<td><inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">T</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Recall (Sensitivity/True Positive Rate)</td>
<td>Proportion of actual positives correctly identified. High recall implies fewer missed detections.</td>
<td><inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">T</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">N</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Average Precision (AP)</td>
<td>Area under the precision&#x2013;recall curve for a single class <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>q</mml:mi></mml:math></inline-formula>, capturing the trade-off between precision and recall.</td>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:msub><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">R</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mrow><mml:mi mathvariant="normal">R</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Mean Average Precision (mAP)</td>
<td>Mean of AP across all <italic>Q</italic> object classes. In this work (<inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, vehicle detection), <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow></mml:math></inline-formula>.</td>
<td><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Q</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mi>q</mml:mi></mml:msub></mml:mstyle></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<p>All experiments were conducted using the PyTorch implementation of YOLOv8 on a workstation equipped with an NVIDIA RTX 3060 GPU (12 GB VRAM), Intel Core i7 processor, and 32 GB RAM. The custom vehicle dataset contained a total of 177 images (168 for training and 9 for testing), annotated using the YOLO format. The model was trained for 200 epochs with an initial learning rate of 0.001 and batch size of 16. Data augmentation techniques such as random flipping, scaling, and brightness normalization were employed to improve generalization across diverse conditions. Complete training configuration and evaluation metrics are summarized in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Model configuration and training details.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Value/Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Framework</td>
<td>PyTorch with YOLOv8 base configuration</td>
</tr>
<tr>
<td>Learning rate</td>
<td>0.001</td>
</tr>
<tr>
<td>Batch size</td>
<td>16</td>
</tr>
<tr>
<td>Epochs</td>
<td>200</td>
</tr>
<tr>
<td>Optimizer</td>
<td>AdamW</td>
</tr>
<tr>
<td>EMA decay factor</td>
<td><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.9998</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Data augmentation</td>
<td>Mosaic cropping, horizontal flipping, brightness normalization</td>
</tr>
<tr>
<td>Evaluation metrics</td>
<td>Precision, recall, mAP@0.5, mAP@0.5&#x2013;0.95</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Result</title>
<p>To comprehensively evaluate the effectiveness of the proposed EMA-GhostConv YOLOv8, we present both qualitative and quantitative analyses on a custom vehicle detection dataset. The qualitative results illustrate the model&#x2019;s behavior in real-world scenarios, particularly under challenging conditions such as occlusion, low lighting, and scale variation, while the quantitative results provide objective performance metrics that validate the improvements in accuracy, precision, recall, and robustness over the baseline YOLOv8 architecture.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Qualitative Results</title>
<p>To visually assess the detection performance under real-world conditions, we present comparative visualizations in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, which displays side-by-side comparisons of original input frames (left) and their corresponding detection outputs from the proposed EMA-GhostConv YOLOv8 model (right). The figure highlights the model&#x2019;s ability to accurately localize vehicles across diverse scenarios, including heavy traffic, partial occlusions, varying illumination, and multi-scale objects, while maintaining clean bounding box boundaries and minimizing false positives. Unlike the baseline YOLOv8, which often misses occluded or small vehicles or generates redundant detections, the enhanced architecture demonstrates consistent robustness and higher confidence in complex urban environments. These qualitative improvements align with and support the quantitative gains reported in the following section.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Qualitative detection performance of the proposed EMA-GhostConv YOLOv8. <italic>Left side:</italic> Original input images capturing real-world traffic scenes. <italic>Right side:</italic> Corresponding detection outputs generated by the model, with red bounding boxes indicating detected vehicles.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-4.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> illustrates the training dynamics of both the baseline and proposed models over 200 epochs, while <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents the corresponding training and testing loss curves. On both the KITTI and Custom datasets, the proposed EMA-GhostConv YOLOv8 exhibits faster convergence and consistently lower loss values from early epochs onward, indicating more stable and efficient optimization. Concurrently, its precision, recall, mAP@50, and mAP@50&#x2013;95 metrics remain persistently higher than those of the baseline across all epochs. The smoother loss trajectories and elevated plateau levels in the metric curves suggest that the EMA layer effectively stabilizes gradient updates by smoothing feature representations, while GhostConv enhances feature expressiveness without introducing redundancy.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Training convergence curves for precision, recall, mAP@50, and mAP@50&#x2013;95 over 200 epochs on the KITTI and Custom datasets, comparing baseline YOLOv8n (left) and proposed EMA-GhostConv YOLOv8 (right).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Training and testing loss curves of baseline YOLOv8n and proposed EMA-GhostConv YOLOv8 on KITTI and Custom datasets over 200 epochs, showing faster convergence and lower loss for the proposed model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-6.tif"/>
</fig>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, the proposed EMA-GhostConv YOLOv8 consistently outperforms the YOLOv8n baseline on both the KITTI and Custom datasets, achieving substantial improvements in precision (&#x002B;9.63% on KITTI, &#x002B;15.35% on Custom), recall (&#x002B;9.79%, &#x002B;12.69%), and mAP@50 (&#x002B;7.57%, &#x002B;6.60%). Remarkably, these gains are attained with only a marginal reduction in inference speed, demonstrating that the proposed enhancements significantly boost detection accuracy without compromising real-time performance. This favorable accuracy&#x2014;efficiency balance underscores the model&#x2019;s suitability for deployment in latency-sensitive intelligent transportation systems (ITS).</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of the proposed EMA-GhostConv YOLOv8 model with baseline YOLOv8n on KITTI and Custom datasets: (a) computational complexity and (b) detection performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center" colspan="9"><bold>(a) Computational complexity</bold></th>
</tr>
<tr>
<th>Model</th>
<th/>
<th align="center" colspan="3">KITTI Dataset</th>
<th align="center" colspan="3">Custom Dataset</th>
<th/>
</tr>
<tr>
<th></th>
<th></th>
<th>Params (M)</th>
<th>GFLOPs</th>
<th>FPS</th>
<th>Params (M)</th>
<th>GFLOPs</th>
<th>FPS</th>
<th></th>
</tr>
</thead>
<tbody>
<tr>
<td colspan="2">YOLOv8n (Baseline)</td>
<td>3.2</td>
<td>6.7</td>
<td>78.10</td>
<td>3.2</td>
<td>6.7</td>
<td>75.80</td>
<td></td>
</tr>
<tr>
<td colspan="2">EMA-GC YOLOv8 (Proposed)</td>
<td>2.9</td>
<td>6.1</td>
<td>81.45</td>
<td>2.9</td>
<td>6.1</td>
<td>78.21</td>
<td></td>
</tr>
<tr>
<td colspan="2">Improvement (%)</td>
<td>&#x2212;9.38</td>
<td>&#x2212;8.96</td>
<td>&#x002B;6.00</td>
<td>&#x2212;9.38</td>
<td>&#x2212;8.96</td>
<td>&#x002B;6.41</td>
<td></td>
</tr>
<tr>
<td align="center" colspan="9"><bold>(b) Detection performance</bold></td>
</tr>
<tr>
<td><bold>Model</bold></td>
<td align="center" colspan="4"><bold>KITTI Dataset</bold></td>
<td align="center" colspan="4"><bold>Custom Dataset</bold></td>

</tr>
<tr>
<td></td>
<td><bold>Prec.</bold></td>
<td><bold>Rec.</bold></td>
<td><bold>mAP@50</bold></td>
<td><bold>mAP@50&#x2013;90</bold></td>
<td><bold>Prec.</bold></td>
<td><bold>Rec.</bold></td>
<td><bold>mAP@50</bold></td>
<td><bold>mAP@50&#x2013;90</bold></td>
</tr>
<tr>
<td>YOLOv8n (Baseline)</td>
<td>82.45</td>
<td>76.32</td>
<td>74.10</td>
<td>56.25</td>
<td>80.12</td>
<td>74.52</td>
<td>72.80</td>
<td>54.90</td>
</tr>
<tr>
<td>EMA-GC YOLOv8 (Proposed)</td>
<td>92.08</td>
<td>86.11</td>
<td>81.67</td>
<td>63.42</td>
<td>87.21</td>
<td>82.40</td>
<td>78.65</td>
<td>60.58</td>
</tr>
<tr>
<td>Improvement (%)</td>
<td>&#x002B;11.68</td>
<td>&#x002B;12.84</td>
<td>&#x002B;10.20</td>
<td>&#x002B;12.73</td>
<td>&#x002B;8.85</td>
<td>&#x002B;10.60</td>
<td>&#x002B;8.03</td>
<td>&#x002B;10.35</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Quantitative Results</title>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> presents a comparative confusion matrix analysis of the baseline YOLOv8n and the proposed EMA-GhostConv YOLOv8 across both KITTI and Custom datasets. The diagonal elements (true positives) are significantly higher for the proposed model, indicating superior classification accuracy for the &#x201C;Vehicle&#x201D; class. Notably, the number of false positives (background predicted as vehicle) is drastically reduced, for example, from 167 to 75 on KITTI and from 242 to 124 on the Custom dataset, demonstrating that the EMA and GhostConv modules enhance discrimination capability and reduce over-prediction. This improvement directly contributes to the observed gains in precision and mAP reported in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix comparison between the baseline YOLOv8n model and the proposed EMA-GhostConv YOLOv8 model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="JAI_76274-fig-7.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Ablation Study</title>
<p>To further analyze the impact of each modification, an ablation study was performed by progressively introducing EMA and GhostConv components into the YOLOv8 framework. The ablation study in <xref ref-type="table" rid="table-5">Table 5</xref> quantifies the contribution of each proposed component. Integrating the EMA feature layer alone improves mAP@50 by &#x002B;3.73% on KITTI and &#x002B;4.10% on the Custom dataset over the baseline, demonstrating its effectiveness in stabilizing feature learning and reducing noise, particularly beneficial in the Custom dataset with challenging conditions like occlusion and illumination variance. Replacing standard convolutions with GhostConv modules yields even larger gains (&#x002B;5.35% mAP@50 on KITTI, &#x002B;4.45% on Custom), confirming that enhanced feature diversity with reduced redundancy significantly boosts detection capability. Most importantly, the full EMA-GhostConv YOLOv8 configuration achieves the highest performance across all metrics on both datasets, with mAP@50 improvements of &#x002B;7.57% (KITTI) and &#x002B;6.60% (Custom) over the baseline. This consistent and additive improvement confirms that EMA and GhostConv operate synergistically: EMA enhances temporal and spatial feature consistency, while GhostConv enriches representational capacity with minimal computational cost, together delivering both robustness and efficiency.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance comparison of YOLOv8 configurations on KITTI and custom datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th align="center" colspan="3">KITTI Dataset</th>
<th align="center" colspan="3">Custom Dataset</th>
</tr>
<tr>
<th></th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>mAP@50 (%)</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>mAP@50 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Baseline YOLOv8</td>
<td>82.45</td>
<td>76.32</td>
<td>74.10</td>
<td>78.10</td>
<td>74.52</td>
<td>75.80</td>
</tr>
<tr>
<td>YOLOv8 &#x002B; EMA Layer</td>
<td>86.71</td>
<td>80.54</td>
<td>77.83</td>
<td>86.50</td>
<td>81.70</td>
<td>79.90</td>
</tr>
<tr>
<td>YOLOv8 &#x002B; GhostConv</td>
<td>89.64</td>
<td>83.22</td>
<td>79.45</td>
<td>89.00</td>
<td>83.12</td>
<td>80.25</td>
</tr>
<tr>
<td>EMA-GC YOLOv8 (Full)</td>
<td>92.08</td>
<td>86.11</td>
<td>81.67</td>
<td>93.45</td>
<td>87.21</td>
<td>82.40</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Comparison Analysis</title>
<p>To rigorously assess the competitiveness of the proposed framework, we compare the EMA-GhostConv YOLOv8 model against several widely adopted object detectors, including SSD, Faster R-CNN, RetinaNet, DETR, EfficientDet-D2, YOLOv5, YOLOv7, and different YOLOv8 variants, under consistent experimental conditions. As summarized in <xref ref-type="table" rid="table-6">Table 6</xref>, the proposed model achieves the highest scores across all major evaluation metrics on both the KITTI and Custom datasets. Specifically, EMA-GC YOLOv8 attains 81.67% mAP@50 on KITTI and 82.40% on the Custom dataset, surpassing the next-best YOLOv8l and YOLOv7 models by margins of &#x002B;2.22 and &#x002B;2.15 mAP@50 points, respectively. Furthermore, the proposed approach exhibits strong generalization with consistent improvements in precision and recall, achieving 92.08% precision and 86.11% recall on KITTI, and 93.45% precision and 87.21% recall on the Custom dataset. These results highlight that the integration of the EMA feature enhancement and GhostConv modules substantially improves detection robustness, particularly under challenging scenarios such as partial occlusion, scale variation, and illumination changes, while preserving the real-time efficiency essential for intelligent transportation and autonomous driving applications. The consistent gains across both standard and custom datasets underscore the generalizability and practical viability of the proposed model.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison of the proposed EMA-GhostConv YOLOv8 model with state-of-the-art detectors on KITTI and Custom datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th align="center" colspan="4">KITTI Dataset</th>
<th align="center" colspan="4">Custom Dataset</th>
</tr>
<tr>
<th></th>
<th>Prec.</th>
<th>Rec.</th>
<th>mAP@50</th>
<th>mAP@50&#x2013;90</th>
<th>Prec.</th>
<th>Rec.</th>
<th>mAP@50</th>
<th>mAP@50&#x2013;90</th>
</tr>
</thead>
<tbody>
<tr>
<td>Faster R-CNN</td>
<td>79.24</td>
<td>72.10</td>
<td>70.35</td>
<td>55.42</td>
<td>75.80</td>
<td>70.42</td>
<td>69.10</td>
<td>54.30</td>
</tr>
<tr>
<td>SSD</td>
<td>68.12</td>
<td>63.27</td>
<td>66.41</td>
<td>50.28</td>
<td>70.25</td>
<td>65.11</td>
<td>67.80</td>
<td>52.46</td>
</tr>
<tr>
<td>RetinaNet</td>
<td>86.40</td>
<td>80.25</td>
<td>77.10</td>
<td>60.30</td>
<td>87.20</td>
<td>82.10</td>
<td>79.50</td>
<td>62.80</td>
</tr>
<tr>
<td>DETR</td>
<td>84.70</td>
<td>78.90</td>
<td>75.80</td>
<td>59.10</td>
<td>85.60</td>
<td>80.40</td>
<td>77.90</td>
<td>61.20</td>
</tr>
<tr>
<td>EfficientDet-D2</td>
<td>87.30</td>
<td>81.60</td>
<td>78.40</td>
<td>61.80</td>
<td>88.10</td>
<td>82.90</td>
<td>80.60</td>
<td>64.10</td>
</tr>
<tr>
<td>YOLOv5</td>
<td>85.10</td>
<td>78.93</td>
<td>76.82</td>
<td>60.75</td>
<td>86.50</td>
<td>81.70</td>
<td>79.90</td>
<td>63.10</td>
</tr>
<tr>
<td>YOLOv7</td>
<td>89.34</td>
<td>82.56</td>
<td>78.92</td>
<td>62.84</td>
<td>89.00</td>
<td>83.12</td>
<td>80.25</td>
<td>64.85</td>
</tr>
<tr>
<td><bold>EMA-GC YOLOv8 (Proposed)</bold></td>
<td><bold>92.08</bold></td>
<td><bold>86.11</bold></td>
<td><bold>81.67</bold></td>
<td><bold>65.37</bold></td>
<td><bold>93.45</bold></td>
<td><bold>87.21</bold></td>
<td><bold>82.40</bold></td>
<td><bold>67.42</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion and Future Work</title>
<p>This study presents the EMA-GhostConv YOLOv8, a refined vehicle detection framework that enhances feature stability and fusion efficiency by integrating an Exponential Moving Average (EMA) layer and GhostConv-based modules within the neck structure. Extensive experiments on both the KITTI benchmark and a challenging custom vehicle dataset demonstrate the proposed model&#x2019;s consistent superiority over the YOLOv8n baseline and several state-of-the-art detectors, including YOLOv5, YOLOv7, Faster R-CNN, SSD, RetinaNet, DETR, and EfficientDet-D2. The EMA-GC YOLOv8 achieves the highest precision, recall, and mean Average Precision (mAP) values, as summarized in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>

<p>Ablation studies (<xref ref-type="table" rid="table-5">Table 5</xref>) further validate the individual and synergistic benefits of the EMA and GhostConv components, confirming that these targeted architectural enhancements substantially improve detection robustness under challenging conditions such as partial occlusion, varying illumination, and dense traffic scenes. The proposed model achieves an optimal trade-off between accuracy and computational efficiency, making it highly suitable for deployment in edge-based intelligent transportation systems and real-time surveillance applications.</p>

<p>Future work will explore incorporating temporal feature modeling for video-based object detection, cross-domain adaptation techniques to enhance generalization across diverse environments, and lightweight quantization strategies to further optimize the model for embedded and resource-constrained platforms.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to thank the open-source research community for providing valuable resources that supported the development and evaluation of this work.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>A. S. M. Masudur Rahman led the conceptualization, methodology design, software implementation, formal analysis and original manuscript preparation. Muhammad Zunair Zamir contributed to methodology refinement, conducted the literature review, performed validation and participated in writing and editing. Syed Sajid Ullah supported methodology design, contributed to formal analysis and assisted in original manuscript preparation. Salman Khan provided support in validation and contributed to proofreading. Maria Saman assisted in data curation and contributed to preliminary result verification. Naqash Bahadar supported project administration and participated in reviewing and refining the manuscript. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The KITTI dataset used in this study is publicly available at <ext-link ext-link-type="uri" xlink:href="http://www.cvlibs.net/datasets/kitti/">http://www.cvlibs.net/datasets/kitti/</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Redmon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Divvala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Farhadi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>You only look once: unified, real-time object detection</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>779</fpage>&#x2013;<lpage>88</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Redmon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Farhadi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>YOLO9000: better, faster, stronger</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>7263</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Redmon</surname> <given-names>J</given-names></string-name>, <string-name><surname>Farhadi</surname> <given-names>A</given-names></string-name></person-group>. <article-title>YOLOv3: an incremental improvement</article-title>. <comment>arXiv:1804.02767. 2018</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bochkovskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <article-title>YOLOv4: optimal speed and accuracy of object detection</article-title>. <comment>arXiv:2004.10934. 2020</comment>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jocher</surname> <given-names>G</given-names></string-name></person-group>. <article-title>YOLOv5 by ultralytics</article-title> <comment>2020 [Internet]. [cited 2025 Dec 23]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/ultralytics/yolov5">https://github.com/ultralytics/yolov5</ext-link>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Department</surname> <given-names>MVI</given-names></string-name></person-group>. <article-title>YOLOv6: a single-stage object detection framework for industrial applications</article-title>. <comment>arXiv:2209.02976. 2022</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Bochkovskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <article-title>YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors</article-title>. <comment>arXiv:2207.02696. 2023</comment>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Ultralytics</collab></person-group>. <article-title>YOLOv8: cutting-edge real-time object detection in native Python</article-title>. <comment>arXiv:2402.13616. 2024</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Yeh</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <chapter-title>YOLOv9: learning what you want to learn using programmable gradient information</chapter-title>. In: <source>Computer Vision&#x2014;ECCV 2024 (ECCV 2024)</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>J</given-names></string-name></person-group>. <article-title>YOLOv10: real-time end-to-end object detection</article-title>. <comment>arXiv:2405.14458. 2024</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yadav</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Nigam</surname> <given-names>N</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Choudhary</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Improving vehicle detection using ghost convolution and ghostbottleneck layers in YOLOv5s</article-title>. In: <conf-name>2023 International Conference on Computational Intelligence and Sustainable Engineering Solutions (CISES)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>322</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bao</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>R</given-names></string-name></person-group>. <article-title>YED-YOLO: an object detection algorithm for automatic driving</article-title>. <source>Signal Image Video Process</source>. <year>2024</year>;<volume>18</volume>(<issue>10</issue>):<fpage>7211</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11760-024-03387-8</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Optimizing road safety: lightweight YOLOv8 models and GhostC2f for real-time distracted driving detection</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>21</issue>):<fpage>8844</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s23218844</pub-id>; <pub-id pub-id-type="pmid">37960543</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lv</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>GS-YOLO: a lightweight SAR ship detection model based on enhanced GhostNetV2 and SE attention mechanism</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>108414</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3438797</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Han</surname> <given-names>K</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>GhostNetv2: Enhance cheap operation with long-range attention</chapter-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<issue>35</issue>:<fpage>9969</fpage>&#x2013;<lpage>82</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Enhancing real-time vehicle detection with a lightweight model</article-title>. <source>J Intell Fuzzy Syst</source>. <year>2024</year>;<volume>49</volume>(<issue>6</issue>):<fpage>1452</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1177/18758967251348957</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ullah</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Zamir</surname> <given-names>MZ</given-names></string-name>, <string-name><surname>Ishfaq</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Attention-augmented YOLOv8 with ghost convolution for real-time vehicle detection in intelligent transportation systems</article-title>. <source>J Artif Intell</source>. <year>2025</year>;<volume>7</volume>:<fpage>255</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.32604/jai.2025.069008</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A commodity object detection method based on EMA-YOLO</article-title>. In: <conf-name>2025 10th International Conference on Intelligent Computing and Signal Processing (ICSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>457</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Han</surname> <given-names>X</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Lightweight ship object detection based on CA and EMA</article-title>. In: <conf-name>2024 9th International Conference on Automation, Control and Robotics Engineering (CACRE)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>231</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Lightweight real-time road defect detection with multi-coordinate attention</article-title>. <source>J Real Time Image Process</source>. <year>2025</year>;<volume>22</volume>(<issue>5</issue>):<fpage>168</fpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bakirci</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Advanced aerial monitoring and vehicle classification for intelligent transportation systems with YOLOv8 variants</article-title>. <source>J Netw Comput Appl</source>. <year>2025</year>;<volume>237</volume>:<fpage>104134</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jnca.2025.104134</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>H</given-names></string-name></person-group>. <article-title>YOLOv8-PD: an improved road damage detection algorithm based on YOLOv8n model</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>12052</fpage>. doi:<pub-id pub-id-type="doi">10.21203/rs.3.rs-4199735/v1</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Multi-object detection for daily road maintenance inspection with UAV based on improved YOLOv8</article-title>. <source>IEEE Trans Intell Transp Syst</source>. <year>2024</year>;<volume>25</volume>(<issue>11</issue>):<fpage>16548</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tits.2024.3437770</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Elhenidy</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Labib</surname> <given-names>LM</given-names></string-name>, <string-name><surname>Haikal</surname> <given-names>AY</given-names></string-name>, <string-name><surname>Saafan</surname> <given-names>MM</given-names></string-name></person-group>. <article-title>GY-YOLO: ghost separable YOLO for pedestrian detection</article-title>. <source>Neural Comput Appl</source>. <year>2025</year>;<volume>37</volume>:<fpage>14907</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-025-11207-4</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>SED-YOLO based multi-scale attention for small object detection in remote sensing</article-title>. <source>Sci Rep</source>. <year>2025</year>;<volume>15</volume>(<issue>1</issue>):<fpage>3125</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-025-87199-x</pub-id>; <pub-id pub-id-type="pmid">39856170</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname></string-name> <string-name> <given-names>Z</given-names></string-name></person-group>. <article-title>Mp-YOLO: multidimensional feature fusion based layer adaptive pruning yolo for dense vehicle object detection algorithm</article-title>. <source>J Vis Commun Image Rep</source>. <year>2025</year>;<volume>112</volume>:<fpage>104560</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jvcir.2025.104560</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Thatikonda</surname> <given-names>M</given-names></string-name></person-group>. <article-title>An enhanced real-time object detection of helmets and license plates using a lightweight YOLOv8 deep learning model [master&#x2019;s thesis]. Dayton, OH, USA: Wright State University</article-title>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ouyang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>The vehicle object detection algorithm based on improved YOLOv8</article-title>. In: <conf-name>International Conference on Artificial Intelligence and Autonomous Transportation</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>X</given-names></string-name></person-group>. <article-title>LVD-YOLO: an efficient lightweight vehicle detection model for intelligent transportation systems</article-title>. <source>Image Vis Comput</source>. <year>2024</year>;<volume>151</volume>:<fpage>105276</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.imavis.2024.105276</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A multi-objective dynamic detection model in autonomous driving based on an improved YOLOv8</article-title>. <source>Alex Eng J</source>. <year>2025</year>;<volume>122</volume>:<fpage>453</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aej.2025.03.020</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>