<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">74911</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.074911</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Intelligent Orchard Anti-Damage System Combining Real-Time AI Image Recognition and Laser-Based Deterrence for Multi-Target Monkeys</article-title>
<alt-title alt-title-type="left-running-head">An Intelligent Orchard Anti-Damage System Combining Real-Time AI Image Recognition and Laser-Based Deterrence for Multi-Target Monkeys</alt-title>
<alt-title alt-title-type="right-running-head">An Intelligent Orchard Anti-Damage System Combining Real-Time AI Image Recognition and Laser-Based Deterrence for Multi-Target Monkeys</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Cho</surname><given-names>Shih-Ming</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Sung-Wen</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Chiu</surname><given-names>Min-Chie</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>mcchiu@gm.ttu.edu.tw</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Shao-Chun</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science and Engineering, Tatung University, No.40, Sec. 3, Zhongshan N. Rd.</institution>, <addr-line>Taipei City, 10452</addr-line>, <country>Taiwan</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Mechanical and Materials Engineering, Tatung University</institution>, <addr-line>No. 40, Sec. 3, Zhongshan N. Rd., Taipei City, 10452</addr-line>, <country>Taiwan</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Min-Chie Chiu. Email: <email>mcchiu@gm.ttu.edu.tw</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>12</day><month>3</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>2</issue>
<elocation-id>37</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>12</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_74911.pdf"></self-uri>
<abstract>
<p>To address crop depredation by intelligent species (e.t, macaques) and the habituation from traditional methods, this study proposes an intelligent, closed-loop, adaptive laser deterrence system. A core contribution is an efficient multi-stage Semi-Supervised Learning (SSL) and incremental fine-tuning (IFT) framework, which reduced manual annotation by &#x007E;60% and training time by &#x007E;68%. This framework was benchmarked against YOLOv8n, v10n, and v11n. Our analysis revealed that YOLOv12n&#x2019;s high Signal-to-Noise Ratio (SNR) (47.1% retention) pseudo-labels made it the only model to gain performance (&#x002B;0.010 mAP) from SSL, allowing it to overtake competitors. Subsequently, in the IFT stress test, YOLOv12n proved most robust (a minimal &#x2212;0.019 mAP decline), whereas YOLOv10n suffered catastrophic failure (&#x2212;0.233 mAP), highlighting its incompatibility with IFT. The final model achieved high performance (mAP@0.5 of 0.947 for macaques, 0.946 for laser spots). In Multi-Object Tracking (MOT), this study quantitatively confirms that Bottom-Up Tracking by Sorting (BoT-SORT) (1.88 s avg. tracklet lifetime) significantly outperforms ByteTrack (0.81 s) in identity preservation for visually similar macaques. System integration achieved 480 Frames Per Second (FPS) real-time inference on edge devices. A quadratic polynomial fitting model ensured high-precision aiming (RMSE &#x003C; 2 pixels; best 1.2 pixels) by compensating for distortion. To fundamentally solve habituation, an adaptive strategy driven by a Deep Deterministic Policy Gradient (DDPG) framework was introduced. By using a habituation penalty term (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) to force unpredictable sequences, the DDPG strategy achieved a stable 88% average Intrusion Frequency Reduction Rate (IFRR) in field experiments, suppressing habituation in highly intelligent species. This study develops an efficient, precise, low-cost, and habituation-resistant automated wildlife defense system.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multi-target tracking</kwd>
<kwd>artificial intelligence recognition</kwd>
<kwd>laser calibration</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<sec id="s1_1">
<label>1.1</label>
<title>Motivation and Objectives</title>
<p>Crop depredation by wildlife poses a critical threat to agricultural sustainability. While traditional methods utilizing infrared sensors [<xref ref-type="bibr" rid="ref-1">1</xref>] or passive barriers may suffice for avian pests [<xref ref-type="bibr" rid="ref-2">2</xref>], they prove largely ineffective against highly intelligent non-human primates, particularly the Formosan macaque (<italic>Macaca cyclopis</italic>). In Taiwan, the conflict is intensified by legal protections and habitat degradation, resulting in a population surge to approximately 300,000 and frequent agricultural intrusion. However, current mitigation strategies face an insurmountable biological bottleneck: Cognitive Adaptability. As noted by Koirala et al. [<xref ref-type="bibr" rid="ref-3">3</xref>], primates rapidly habituate to predictable deterrent patterns&#x2014;such as periodic noise or fixed laser scans. Furthermore, traditional approaches as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> suffer from structural limitations: passive electric fences are cost-prohibitive to maintain, while active traps and Passive Infrared Sensor (PIR) are plagued by low efficiency and the risk of capturing non-target species [<xref ref-type="bibr" rid="ref-4">4</xref>]. Consequently, the development of an intelligent, closed-loop system capable of generating unpredictable, adaptive stimuli is not merely an enhancement but a necessity for resolving such advanced human-wildlife conflicts.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Traditional wildlife deterrence strategies and their limitations</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-1.tif"/>
</fig>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>System Innovations and Core Contributions</title>
<p>To address the significant challenges of crop depredation by highly intelligent wildlife (e.g., macaques) and the habituation associated with traditional deterrent methods, this study proposes an intelligent, closed-loop identification and deterrence system equipped with adaptive strategies. This system integrates real-time identification, precision calibration, and effective countermeasures. Its core technological breakthroughs and methodological contributions are summarized as follows:
<list list-type="bullet">
<list-item>
<p>A High-Performance and High-Robustness Artificial Intelligence (AI) Training Framework: We propose a multi-stage training framework integrating Semi-Supervised Learning (SSL) and Incremental Fine-Tuning (IFT). This methodology not only drastically reduces the reliance on manual annotation and training time but also theoretically solves the &#x201C;Catastrophic Forgetting&#x201D; problem when the model learns new functionalities (such as laser spots).</p></list-item>
<list-item>
<p>A Highly Discriminative Multi-Object Tracking Methodology: Addressing the core challenge of &#x201C;high intra-class similarity&#x201D; (similar appearance) within macaque groups, this study empirically validates and establishes that the Bot-SORT tracker, by leveraging its highly discriminative Re-ID features, is significantly more robust in identity preservation than traditional methods (like ByteTrack).</p></list-item>
<list-item>
<p>Adaptive Decision-Making (Anti-Habituation) and High-Precision Execution: This system is the first to integrate a Deep Deterministic Policy Gradient (DDPG) reinforcement learning framework as its decision core. By penalizing fixed patterns in the reward function (i.e., rewarding entropy), the DDPG drives the system to generate unpredictable laser strategies, fundamentally suppressing the habituation effect in highly intelligent species. This strategy is enabled by a quadratic polynomial fitting model for dynamic laser calibration, ensuring high-precision targeting.</p></list-item>
</list></p>
<p>These contributions provide a solid methodological foundation for developing high-precision, automated wildlife defense systems capable of overcoming the challenge of habituation.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Research Methodology Evaluation</title>
<sec id="s2_1">
<label>2.1</label>
<title>Limitations of Traditional Deterrence and Early Detection Systems</title>
<p>Human-wildlife conflict (HWC) has emerged as a complex global issue, threatening both food security and local livelihoods [<xref ref-type="bibr" rid="ref-5">5</xref>]. However, a fundamental limitation of traditional deterrents (e.g., noise, fences) is their predictability, which allows highly intelligent species like macaques to quickly habituate to them, rendering these methods ineffective [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>To overcome this challenge, &#x201C;Intelligent Deterrence&#x201D; has emerged, centered on implementing an Anti-Habituation strategy. Based on the generalized analysis of existing literature, we categorize current research into different levels of complexity (Level of Anti-Habituation, LoAH):
<list list-type="bullet">
<list-item>
<p>LoAH 1&#x2013;2 (Basic AI): Systems advanced to using AI for species identification to trigger &#x201C;species-specific&#x201D; deterrents (e.g., predator calls) [<xref ref-type="bibr" rid="ref-7">7</xref>], an improvement over randomization (LoAH 1).</p></list-item>
<list-item>
<p>LoAH 3 (Multi-Modal): To counter habituation to single stimuli, systems evolved to &#x201C;multi-modal escalation.&#x201D; As exemplified by Park and Shim [<xref ref-type="bibr" rid="ref-8">8</xref>], AI is used to target specific animals (e.g., wild boars) and trigger multiple sensory deterrents (e.g., light and sound).</p></list-item>
<list-item>
<p>LoAH 4 (Mobility): This level recognizes that a &#x201C;fixed location&#x201D; is itself a source of habituation. It introduces &#x201C;mobility,&#x201D; using UAVs to mimic aerial predators [<xref ref-type="bibr" rid="ref-9">9</xref>] or UGVs [<xref ref-type="bibr" rid="ref-10">10</xref>] to mimic the patrolling behavior of terrestrial predators (e.g., coyotes). However, these systems often face significant challenges regarding limited endurance or high initial costs.</p></list-item>
<list-item>
<p>LoAH 5 (Adaptive): This is the final frontier of deterrence. As conceptualized by Mishra and Yadav [<xref ref-type="bibr" rid="ref-11">11</xref>], this &#x201C;closed-loop&#x201D; system features an AI that is not just a detector, but a learner. It observes the animal&#x2019;s reaction and &#x201C;proactively changes its strategy&#x201D; if habituation is detected, maintaining long-term efficacy.</p></list-item>
</list></p>
<p>In summary, the literature demonstrates a clear progression from static, predictable deterrence (LoAH 1&#x2013;2) toward dynamic, unpredictable strategies (LoAH 4&#x2013;5). However, while LoAH 4 (Mobility) suffers from high deployment costs, LoAH 5 (Adaptive) remains the conceptual ideal, lacking concrete, low-cost, and empirically validated solutions. The system proposed in this study combines the cost-effectiveness of a static, edge-deployed platform with the highest level of anti-habituation capabilities (Level 5), endowed by reinforcement learning. We are the first to introduce Deep Reinforcement Learning (DDPG) into the decision-making core of laser deterrence, enabling the system to achieve closed-loop deterrence via an Adaptive Feedback Loop (as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>). By penalizing fixed patterns within the reward function, our system autonomously learns and generates unpredictable laser repulsion sequences, fundamentally overcoming the habituation risk inherent in static systems. To concretely quantify these advantages, we provide a systematic comparison of our system against representative intelligent deterrence systems from the literature in <xref ref-type="table" rid="table-1">Table 1</xref> below, highlighting the key differences in anti-habituation strategies, deployment costs, and quantified efficacy.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Dual-Layer, cognitive closed-loop wildlife deterrence system based on &#x201C;perception-decision-action&#x201D; framework</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-2.tif"/>
</fig><table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparative analysis of intelligent wildlife deterrence systems</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Comparison aspect</th>
<th>This System (AI &#x002B; RL &#x002B; Laser)</th>
<th>SARD (Mishra, 2024)</th>
<th>UAV (Afridi, 2025)</th>
<th>UGV (Aurora Project)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Deployment platform</td>
<td>Static (Edge AI)</td>
<td>Static (Edge AI)</td>
<td>Aerial (UAV)</td>
<td>Ground (UGV)</td>
</tr>
<tr>
<td>Core AI<bold>/</bold>Sensor</td>
<td>YOLOv12n &#x002B; BoT-SORT &#x002B; DDPG RL</td>
<td>R-CNN/PIR</td>
<td>(Autonomous)/Vision</td>
<td>(Autonomous)/Vision</td>
</tr>
<tr>
<td>Deterrence mechanism</td>
<td>Precision Green Laser</td>
<td>Species-specific Ultrasound</td>
<td>Physical Presence/Noise</td>
<td>Predator Mimicry/Gas Cannon</td>
</tr>
<tr>
<td>Anti-habituation strategy</td>
<td>Adaptive Feedback (LoAH 5)<break/>DDPG Dynamic Adjustment</td>
<td>Adaptive (LoAH 5)<break/>AI-based decision trigger</td>
<td>Mobility (LoAH 4)</td>
<td>Mobility &#x002B; Biomimicry (LoAH 4)</td>
</tr>
<tr>
<td>Quantified efficacy</td>
<td>88% reduction in intrusion frequency</td>
<td>92% removal success rate, 30% reduction in crop loss</td>
<td>Clears 8 hectares in<break/> 10 min</td>
<td>In testing, no displacement data provided</td>
</tr>
<tr>
<td>Cost/Key limitations</td>
<td>Low cost<break/> requires DDPG learning time</td>
<td>Low cost<break/> limited adaptive strategy (non-RL), potential for long-term habituation</td>
<td>Very short endurance<break/> high labor/regulatory costs</td>
<td>Extremely high initial cost ($70,000 USD)<break/> terrain-limited</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Comparison of Deep Learning-Based Object Detection Techniques</title>
<p>Object detection is a foundational AI computer vision module for real-time monitoring applications, and the YOLO (You Only Look Once) series has become the mainstream single-stage detector due to its superior balance of speed and accuracy. Yang et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] employed an improved YOLOv5s model for real-time wildlife detection encompassing multiple animal targets, including macaques (<italic>Macaca mulatta</italic>), and achieved high accuracy and real-time performance in forest environments; Wang et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] further indicated that a YOLOv5-m trained on mixed day- and night-time images can achieve an mAP &#x003D; 0.72. However, Liu et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] emphasized that adverse conditions such as low light and fog remain a severe challenge. To address the demand for dynamic tracking, Yang et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] adopted YOLOv7 combined with DeepSORT to achieve real-time identity tracking of wildlife. Wu et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] also validated that YOLOv7 significantly improved detection rates in cluttered environments with vegetation occlusion and lighting variations. Chappidi and Sundaram [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed YOLOv8, which (at 97%) achieved higher accuracy, superior to YOLOv7 (approx. 93%).</p>
<p>This study adopts YOLOv12 (a model implemented by the Ultralytics framework) as the backbone network, using its YOLOv12n pre-trained weights as the starting point for fine-tuning. Its attention mechanism enhances recognition capability for small targets and complex backgrounds while maintaining real-time inference speed. YOLOv12n exhibits excellent computational efficiency; Ref. [<xref ref-type="bibr" rid="ref-18">18</xref>] reports its inference latency on a T4 GPU is only 1.64 ms, and its COCO mAP is 2.1% higher than YOLOv10n. This balance of speed and accuracy makes it highly suitable for resource-constrained edge computing applications.</p>
<p>As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, this study develops a system where the YOLOv12 model is deployed on an edge AI computing unit (e.g., NVIDIA Jetson), responsible for real-time image processing. Its detection results (macaques and laser spots) are passed to the laser control module and collaborate with a cloud platform, forming an intelligent closed-loop system.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Schematic diagram of real-time orchard monitoring and repulsion System</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-3.tif"/>
</fig>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Development and Comparative Analysis of Multi-Object Tracking (MOT) for Wildlife Monitoring</title>
<p>With the intensive deployment of camera traps, unmanned aerial vehicles (UAVs), and fixed image-capture platforms in ecological research, large volumes of consecutive imagery are processed across domains to derive information on animal occurrence frequency, activity ranges, and social interactions. In this context, Ogawa et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] noted that the &#x201C;tracking-by-detection&#x201D; (TBD) paradigm has become the mainstream workflow: an object detector first locates targets, followed by a data-association algorithm to generate trajectories across frames. However, the performance of multi-object tracking (MOT) systems remains constrained by two key factors in natural environments: detection accuracy and the robustness of association strategies under conditions of vegetation occlusion, highly similar individual appearances, and abrupt motion.</p>
<p>In recent years, MOT techniques have progressed rapidly, with algorithmic performance continuing to advance across major publicly available benchmarks. This progress largely builds on the seminal work of Wojke et al. [<xref ref-type="bibr" rid="ref-20">20</xref>], who introduced DeepSORT&#x2014;a framework that pioneered the integration of deep-learning appearance features into the tracking pipeline. Subsequent studies have focused on addressing increasingly complex real-world challenges. For example, to mitigate overlooked low-score detections caused by occlusion, Zhang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] developed ByteTrack, employing a novel two-stage matching strategy that effectively utilizes these latent target cues. Likewise, focusing on occlusion scenarios, Cao et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed OC-SORT, which further enhances tracking stability. Building on these advances, Aharon et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] introduced BoT-SORT&#x2014;a comprehensive solution that improves upon the association strategy of ByteTrack by integrating appearance features from Deep-SORT and camera-motion compensation, thereby substantially enhancing tracking robustness in complex scenes. A comparative analysis of the technical features of mainstream MOT algorithms for wildlife monitoring is shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Technical characteristics comparison of mainstream MOT algorithms in wildlife monitoring</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Evaluation aspect</th>
<th>Deep-SORT</th>
<th>ByteTrack</th>
<th>OC-SORT</th>
<th>BoT-SORT</th>
</tr>
</thead>
<tbody>
<tr>
<td>Core principle</td>
<td>Motion &#x002B; Appearance (Re-ID)</td>
<td>Low-score Det. Assoc.</td>
<td>Improved motion model</td>
<td>Motion &#x002B; Appearance Integ.</td>
</tr>
<tr>
<td>Short-term occlusion</td>
<td>Relies on Re-ID</td>
<td>Very Strong (via low-score)</td>
<td>Strong (via motion model)</td>
<td>Very strong (2-stage match)</td>
</tr>
<tr>
<td>Long-term ID (IDF1)</td>
<td>Medium (Re-ID dependent)</td>
<td>Weak (No appearance)</td>
<td>Weak (No appearance)</td>
<td>High (Balances MOTA/IDF1)</td>
</tr>
<tr>
<td>Irregular motion</td>
<td>Weak (KF sensitive to non-linear)</td>
<td>Medium (Needs high FPS/IoU)</td>
<td>Strong (Designed for non-linear)</td>
<td>Strong (Improved KF &#x002B; CMC)</td>
</tr>
<tr>
<td>Cost<bold>/</bold>Speed</td>
<td>High (Extra Re-ID)</td>
<td>Very low (Very fast)</td>
<td>Very Low (Like ByteTrack)</td>
<td>High (Integrates Re-ID)</td>
</tr>
<tr>
<td>Implementation</td>
<td>High (Needs custom Re-ID)</td>
<td>Low (Simple logic)</td>
<td>Low (Pure algorithm)</td>
<td>High (Needs custom Re-ID)</td>
</tr>
<tr>
<td>Wildlife application</td>
<td>Long-term/Cross-cam Re-ID (Needs custom model)</td>
<td>High real-time, short occlusions</td>
<td>Variable animal motion</td>
<td>Fine-grained analysis (Needs max ID stability)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Given the increasing demand for precise, non-invasive management in human-wildlife conflicts, a robust Multi-Object Tracking (MOT) system is essential. While MOT is mature in structured environments, its application in dynamic, ecologically complex scenarios like agriculture presents severe challenges. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, although traditional MOT frameworks integrate detection and tracking, translating them from concept to practical deployment requires overcoming three key challenges identified by this study:</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Schematic diagram of the multi-object tracking (MOT) system</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-4.tif"/>
</fig>
<p><list list-type="bullet">
<list-item>
<p>Unstructured Environments: Agricultural landscapes are inherently heterogeneous (e.g., irregular terrain, dense vegetation, variable lighting), introducing significant noise that can degrade detector and motion predictor performance.</p></list-item>
<list-item>
<p>Visual Homogeneity and Intra-Class Similarity: Macaque groups exhibit high visual similarity among individuals, posing a key challenge to the Re-ID component. This can lead to frequent ID switches and trajectory fragmentation.</p></list-item>
<list-item>
<p>Frequent Occlusions: Macaques are often occluded by vegetation, structures, or other individuals. This disrupts visual continuity, making it difficult for trackers to maintain consistent identities and leading to track loss or false trajectories.</p></list-item>
</list></p>
<p>In summary, this study aims to develop an integrated real-time detection and tracking system specifically optimized for the group behavior of macaques. The experimental focus is to validate this framework&#x2019;s practical efficacy in complex agricultural environments, thereby laying the critical technical foundation for subsequent precise, non-lethal, and intelligent deterrence strategies and providing an innovative solution to the human-macaque conflict.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Foundations of Wildlife Behavior and Evolutionary Assessment of Deterrence Technologies</title>
<p>To overcome the habituation problem of traditional deterrence methods (such as sounds or fixed fences), contemporary research has gradually shifted towards developing intelligent deterrence systems characterized by &#x201C;unpredictability&#x201D; and &#x201C;target specificity&#x201D;. In this context, laser deterrence technology has emerged and been proven to be an effective non-lethal intervention.</p>
<p>In the domain of Green Security Games (GSG), recent advances in adversarial modeling characterize the interaction between defenders and intelligent biological agents (e.g., <italic>Macaca cyclopis</italic>) as a dynamic pursuit-evasion game. Wang et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] demonstrate that static defense policies are inherently vulnerable to &#x201C;best-responding&#x201D; adversaries. These agents rapidly infer deterministic patterns through repeated interactions, effectively optimizing their intrusion policies. This computational perspective frames &#x201C;habituation&#x201D; not merely as biological desensitization, but as adversarial policy exploitation, necessitating the use of stochastic, non-stationary deterrence strategies. To counter this, our system incorporates an adaptive deterrence framework designed to generate unpredictable sequences of stimuli. This framework integrates a parameterized &#x201C;deterrence action space&#x201D;, including variables such as laser flicker frequency, scanning patterns, and multimodal stimuli. The system evaluates the efficacy of each policy based on real-time state feedback from the target (e.g., flight vector or avoidance latency) and dynamically optimizes deterrence actions via the Reinforcement Learning agent.</p>
<p>The application of high-intensity visual stimuli (e.g., lasers) to primates is grounded in robust ethological principles. Recent comparative research by Luongo et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] demonstrates that rodents (e.g., mice) and primates employ fundamentally different visual scene segmentation strategies. In contrast to rodents, primates rely predominantly on the visual system to effectively interact with objects in their environment, thereby supporting the validity of visual intervention strategies as a primary deterrence mechanism.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Research Methodology</title>
<p>This study proposes an intelligent closed-loop deterrence system built upon the &#x201C;Perception-Decision-Action&#x201D; (PDA) theoretical framework (as shown in the system architecture, <xref ref-type="fig" rid="fig-5">Fig. 5</xref>). The system employs a dual-layer collaborative architecture: high-performance Edge AI is responsible for real-time field responses, while the Cloud Platform handles long-term model optimization and data governance.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>System architecture diagram</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-5.tif"/>
</fig>
<p>The system&#x2019;s processing pipeline follows this framework: The Perception Layer (ffmpeg) captures image frames and passes them to the Decision Layer. The AI model in the Decision Layer not only locates &#x201C;macaques&#x201D; but also, critically, identifies the &#x201C;laser spot&#x201D;, which serves as the visual feedback signal for closed-loop control. Once the target is confirmed, the Action Layer&#x2019;s calibration model precisely maps 2D pixel coordinates to 1D physical rotational angles, driving the laser. Concurrently, the Human Machine Interface (HMI) subsystem records all events and provides a &#x201C;human-in-the-loop&#x201D; interface, allowing users to perform strategy adjustments and optimizations.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Perception Subsystem: Multi-Modal Detection and Tracking</title>
<p>The perception subsystem is responsible for accurately detecting and tracking wildlife under diverse and challenging environmental conditions. Its design prioritizes robustness and high precision.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Real-Time Multi-Object Detection and Training Strategy</title>
<p>This detection module utilizes an optimized YOLOv12 architecture, refined through an integrated multi-stage pipeline designed for precision, efficiency, and scalability.</p>
<p>[Stage 1: Foundation Model Construction and High-Precision Teacher Model]. In the foundation model construction phase (Stage 1), we implemented a &#x201C;Target-Exclusive&#x201D; annotation strategy. This approach involved exclusively labeling the target species (macaques) while deliberately retaining distinct but confounding entities&#x2014;such as humans, cats, and dogs&#x2014;as unlabeled data, effectively treating them as &#x201C;implicit background samples&#x201D;. Fundamentally, this strategy functions as a Semantic High-Pass Filter: by compelling the model to penalize feature activations associated with non-target entities as errors, the network is forced to sharpen its decision boundaries. This mechanism fundamentally suppresses false positives arising from morphologically similar biological features. Furthermore, the incorporation of approximately 10% pure background images into the training set further bolstered the model&#x2019;s noise immunity. Consequently, this mechanism yielded a high Precision of 0.924. Crucially, this metric serves not merely as an indicator of immediate detection performance but as the direct determinant of the Signal-to-Noise Ratio (SNR) for the pseudo-labels generated in the subsequent Semi-Supervised Learning phase (Stage 2). It was empirically verified that the error rate of this teacher model in high-confidence predictions converged to 0.0%, thereby providing a near noise-free supervision signal for Stage 2 and establishing the foundational causal factor that ensures the stability of the semi-supervised learning process.</p>
<p>[Stage 2: SSL &#x0026; Robustness Enhancement]. To scale data diversity efficiently, a Semi-Supervised Learning (SSL) framework is deployed. The S1 teacher model creates pseudo-labels for unlabeled data, leveraging its discriminative power and a strict confidence threshold (<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>0.80</mml:mn></mml:math></inline-formula>) to filter for quality. This ensures a maximal Signal-to-Noise Ratio (SNR) in the training set for the &#x201C;student model&#x201D;. Consequently, the injection of noise-free pseudo-labels optimizes decision boundaries and repairs the recall weaknesses observed in S1, resulting in state-of-the-art performance.</p>
<p>[Stage 3: IFT &#x0026; Functional Expansion]. The final stage employs Incremental Fine-Tuning (IFT) to incorporate the &#x201C;Laser Dot&#x201D; class without triggering &#x201C;Catastrophic Forgetting&#x201D;. IFT is engineered to preserve the perceptual integrity established in SSL, based on the following theoretical principles:
<list list-type="bullet">
<list-item>
<p>Hierarchical Features: Leveraging the distinction between shallow (general) and deep (semantic) feature extraction.</p></list-item>
<list-item>
<p>Parameter Freezing: A hard constraint is applied by freezing the initial backbone layers (<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>f</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>e</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula>), zeroing gradients to lock in prior knowledge.</p></list-item>
<list-item>
<p>Regularized Fine-Tuning: Updates are restricted to the Detection Head using extremely low learning rates (e.g., <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>l</mml:mi><mml:mi>r</mml:mi><mml:mn>0</mml:mn><mml:mo>=</mml:mo><mml:mn>0.0002</mml:mn></mml:math></inline-formula> or <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mn>0.0005</mml:mn></mml:math></inline-formula>) and abbreviated training epochs (40 or 75). This acts as strong L2 regularization, minimizing the disruption of the shared parameter space by the new task.</p></list-item>
</list></p>
<p>By reconciling functional expansion with stability, this multi-stage approach delivers a robust perception subsystem. It provides reliable state space inputs&#x2014;specifically target trajectory and velocity&#x2014;to the Deep Deterministic Policy Gradient (DDPG) decision-making subsystem, serving as the cornerstone for the system&#x2019;s adaptive anti-habituation capabilities.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Multi-Modal Sensor Fusion for Environmental Adaptability</title>
<p>To overcome the limitations of a single sensing modality under adverse weather (rain, fog) and variable lighting (low-light, backlight), the system integrates RGB, Near-Infrared (NIR), and Time-of-Flight (ToF) cameras. Data streams from these sensors are fused and processed with image enhancement algorithms, such as automatic exposure compensation and dehazing, ensuring stable and reliable all-weather detection performance.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Robust Multi-Object Tracking (MOT)</title>
<p>To maintain individual animal identity continuity (crucial for behavioral analysis), this study employs advanced Multi-Object Tracking (MOT) algorithms. The system is designed to compare ByteTrack and BoT-SORT to address rapid movement and mutual occlusion in natural environments. ByteTrack excels at handling large numbers of fast-moving and briefly occluded targets, while BoT-SORT, by integrating appearance features, provides stronger robustness in scenarios involving long-term occlusion and appearance ambiguity. This dual-track strategy ensures the system can establish stable and coherent spatiotemporal trajectories.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Decision Subsystem: Optimization Strategy Based on Reinforcement Learning</title>
<p>This subsystem is the core of the closed-loop system, utilizing a Deep Deterministic Policy Gradient (DDPG) Reinforcement Learning agent to learn the optimal deterrence strategy.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Theoretical Basis and Ethological Rationale</title>
<p>Non-Lethal Deterrence Principle: The system employs a Class 2 safety laser module (output <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo>&#x2264;</mml:mo><mml:mn>5</mml:mn><mml:mrow><mml:mtext>&#xA0;mW</mml:mtext></mml:mrow></mml:math></inline-formula>) compliant with the IEC 60825-1 international standard. This mechanism utilizes the strong contrast between the laser beam and the low-light environment during early morning hours to trigger an avoidance response in macaques, ensuring safe and effective physical intervention.</p>
<p>Anti-Habituation for Highly Intelligent Species: To prevent the rapid habituation caused by traditional methods (e.g., sound, firecrackers) in highly intelligent species like Formosan macaques, the system combines AI tracking with RL decision-making. The goal is to generate unpredictable and target-specific stimulus sequences to maintain long-term deterrent effects.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Hybrid Decision Framework and DDPG Strategy Optimization</title>
<p>The robustness of the DDPG decision policy is strictly predicated on the temporal coherence of the input state vector (<inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>). The reinforcement learning agent relies on accurate history-dependent variables&#x2014;specifically target velocity and dwell time&#x2014;to compute the <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> penalty and optimize the anti-habituation strategy. The superior tracking stability of Bot-SORT (avg. tracklet lifetime 1.88 s) serves as a critical noise-suppression mechanism. By minimizing Identity Switches, Bot-SORT prevents the fragmentation of the state history, ensuring that the DDPG agent receives a high-fidelity representation of the macaque&#x2019;s behavioral response. This causal dependency explains why simpler trackers like ByteTrack, despite acceptable detection accuracy, fail to support the long-term strategic learning required for habituation suppression.</p>
<p>We adopt a hybrid framework combining Reinforcement Learning (RL) and Bayesian Decision Theory. The system utilizes a Deep Deterministic Policy Gradient (DDPG) agent for learning.</p>
<p><bold>State and Action Space:</bold>
<list list-type="bullet">
<list-item>
<p>State Space: Defined by the target attributes provided by the perception subsystem, including the target&#x2019;s real-time trajectory and velocity, distance to the orchard boundary, and target group density.</p></list-item>
<list-item>
<p>Action Space: Corresponds to the laser module&#x2019;s control parameters, including horizontal/vertical angles, power modulation, and deterrence mode (e.g., fixed-point targeting, random scanning, or flicker frequency).</p></list-item>
</list></p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>DDPG Reward Function Design</title>
<p>This study defines the RL agent&#x2019;s optimization goal as a composite reward function (<inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>R</mml:mi></mml:math></inline-formula>). The function balances multiple objectives (long-term efficacy, safety, anti-habituation) through the adjustment of weights (<inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>w</mml:mi></mml:math></inline-formula>).
<list list-type="bullet">
<list-item>
<p>Rationale and Sensitivity Analysis for Weight Setting:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>Primary Consideration (Deterrence Success <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>): The <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is set to the highest value (e.g., 1.0), prioritizing the maximization of protection against crop loss risk (i.e., deterrence success). <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> must dominate the optimization process, ensuring the strategy remains focused on driving macaques away.</p></list-item>
<list-item><label>&#x2013;</label><p>Balance Finding (<inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>): Weights <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (e.g., 0.5 and 0.8) are determined through preliminary parameter Grid Search or Ablation Study in simulation. This process aims to find the optimal balance between deterrence success, energy efficiency, and habituation suppression. If a strategy becomes overly predictable (<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> penalty is insufficient), <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> will be adjusted to increase the Entropy of the action sequence, maintaining long-term deterrent efficacy.</p></list-item>
</list></p></list-item>
</list></p>
<p>Unit and Normalization of the Reward Function: To ensure the mathematical rigor of <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>R</mml:mi></mml:math></inline-formula>, all reward components (<inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) are quantified and normalized before weighting, preventing a single component&#x2019;s magnitude from dominating the optimization. Values from different physical units are converted into Unitless values and scaled to a unified range (e.g., <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> or <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mo>[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>).</p>

<p>Component Definition:
<list list-type="simple">
<list-item><label>&#x2013;</label><p><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Quantified and normalized based on the percentage change in macaque dwell time reduction or exit from the ROI.</p></list-item>
<list-item><label>&#x2013;</label><p><inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Quantified and normalized based on the intermittency of the deterrence action or the degree of compliance with Class 2 laser standards (<inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mspace width="negativethinmathspace" /><mml:mn>5</mml:mn><mml:mrow><mml:mtext>mW</mml:mtext></mml:mrow></mml:math></inline-formula>).</p></list-item>
<list-item><label>&#x2013;</label><p><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Quantified and normalized based on the entropy of the action sequence, ensuring its negative penalty effect is fairly balanced against positive reward terms.</p></list-item>
</list></p>
<p>The structure of the composite reward function is defined as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">d</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">d</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">w</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> listed below presents the DDPG reward function&#x2019;s various components and their design principles.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Reward components and their design principles</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Reward component</th>
<th>Description</th>
<th>Design goal and weight principle</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">d</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Deterrence success</td>
<td>Maximize: Reward target exiting the ROI area immediately after deterrence action or shortening stay time. Weight <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> set to highest (e.g., <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula><sub>r</sub> &#x003D; 1.0).</td>
</tr>
<tr>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">f</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Safety &#x0026; energy efficiency</td>
<td>Maximize: Reward low power (compliant with Class 2 standard, &#x003C;5 mW) and intermittent deterrence actions. Weight <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> set to medium (e.g., <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x003D; 0.5).</td>
</tr>
<tr>
<td><inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">b</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">u</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>Habituation penalty</td>
<td>Minimize (as a penalty term): Penalize the agent for repeating fixed deterrence patterns. Reward the Entropy of the action sequence (unpredictability). Weight <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> used to balance deterrence intensity and long-term effect (e.g., <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x003D; 0.8).</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Through the optimization of this composite function, the DDPG agent can formulate complex strategies in millisecond timeframes, achieving unpredictable and target-specific stimulus sequences, thereby fundamentally suppressing the habituation effect in highly intelligent species.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Actuation Subsystem: High-Precision Calibration and Control</title>
<p>This subsystem constitutes a critical component of the closed-loop control system, responsible for precisely translating commands from the decision engine into physical actions. Its core function is to achieve high-precision, low-latency laser targeting.</p>
<p>To ensure precise laser targeting, the system requires a robust mapping function that converts the target&#x2019;s 2D pixel coordinates on the image plane into a 1D rotation angle in the laser&#x2019;s physical control space. Our system innovatively adopts a <bold>quadratic polynomial fitting model</bold>, rather than a conventional simple linear model, to capture complex non-linear physical characteristics, thereby significantly enhancing calibration accuracy.
<list list-type="bullet">
<list-item>
<p><bold>Dynamic Laser Calibration Model Definition and Mathematical Equation:</bold></p></list-item>
</list></p>
<p>To accurately establish the non-linear mapping relationship between image pixel coordinates (<inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>x</mml:mi></mml:math></inline-formula>) and the laser&#x2019;s physical angle (<inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>y</mml:mi></mml:math></inline-formula>), we employ a Quadratic Polynomial Regression Model. This model assumes a quadratic functional relationship between the horizontal pixel coordinate <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the laser spot in the image and its corresponding horizontal physical rotation angle <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Its mathematical equation is defined as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where:
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the actual physical angle for the <italic>i</italic>-th observation.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the corresponding image pixel coordinate for the <italic>i</italic>-th observation.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the unknown parameters (coefficients) of the model, representing the intercept, linear coefficient, and quadratic coefficient, respectively.</p></list-item>
<list-item>
<p><italic>&#x03F5;</italic><sub><italic>i</italic></sub> is the random error term, representing variance not explained by the model.</p></list-item>
</list></p>
<p>Compared to linear models, which can only fit straight lines, the quadratic polynomial model, by including the <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> term, is capable of fitting a parabola. This enables it to more precisely capture subtle curvilinear relationships caused by lens distortion, minor mechanical mounting errors, or non-linear motor movements. This design choice is expected to significantly reduce calibration errors and provide higher targeting accuracy compared to a linear model.
<list list-type="bullet">
<list-item>
<p><bold>Parameter Estimation and Derivation:</bold></p></list-item>
</list></p>
<p>The model&#x2019;s objective is to find a set of optimal parameter estimates <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that minimize the total error between the model&#x2019;s predicted angle <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and the actually observed angle <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This process is achieved through <bold>Ordinary Least Squares (OLS)</bold>.</p>
<p>The residual is defined as: For each observed data point (<inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), the residual (<inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) is the difference between the actual observed value and the model&#x2019;s predicted value:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The objective function is defined as &#x201C;<bold>Sum of Squared Residuals (SSR)</bold>&#x201D;. The core idea of OLS is to minimize the sum of the squares of the residuals across all observed points. Squaring and summing the residuals prevents positive and negative errors from canceling each other out and gives higher weight to larger errors.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>where <italic>N</italic> is the total number of data points collected during the calibration process.
<list list-type="bullet">
<list-item>
<p><bold>Minimization of SSR and Weight Estimation:</bold></p></list-item>
</list></p>
<p>Our goal is to find the values of <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that minimize the SSR.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula></p>
<p>Mathematically, this is achieved by taking partial derivatives of the SSR with respect to <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, respectively, and setting them equal to zero. Solving this system of simultaneous equations yields the optimal estimates for <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>. The weights estimated through this method, under the assumptions of OLS, are Best Linear Unbiased Estimators (BLUE) according to the Gauss-Markov theorem. We plan to analyze the standard errors of these weight estimates through experimentation to assess their stability and reliability.
<list list-type="bullet">
<list-item>
<p><bold>Practical Application and Performance Verification Methodology</bold></p></list-item>
</list></p>
<p>In the automated calibration procedure, the system will drive the laser to perform a series of angular scans, synchronously recording a pair of (pixel coordinate <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, physical angle <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) data for each position. We will determine the optimal number of calibration points <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>N</mml:mi></mml:math></inline-formula> and data distribution strategy through experimental sensitivity analysis, aiming to achieve the most stable and precise weight estimates. After collecting sufficient data points, the system will automatically invoke the OLS algorithm to calculate the optimal coefficients. Once the model is trained, this computationally inexpensive quadratic equation will be stored. In real-time deterrence tasks, when the AI detects a target at pixel coordinate <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the system can instantaneously (millisecond-level response) calculate the precise angle to which the laser needs to rotate:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mover><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></disp-formula>
<list list-type="bullet">
<list-item>
<p><bold>Performance Verification Methodology:</bold></p></list-item>
</list></p>
<p>To evaluate the actual targeting accuracy of the model, we will design a series of independent real-time targeting tests. These tests will involve randomly placing multiple targets and recording the pixel error between the laser&#x2019;s actual hit point and the AI-detected target center. Ultimately, we will calculate the Root Mean Squared Error (RMSE), serving as a key metric for measuring the system&#x2019;s targeting accuracy, and compare it against the predefined high-precision targeting goal (RMSE &#x003C; 2 pixels). Concurrently, we will conduct weight perturbation analysis and calibration data quality analysis to further validate the scientific basis of the model structure and the robustness of the estimated parameters.</p>
<p>The safety of wildlife and personnel is our top priority. This system employs a Class 2 laser module that complies with the IEC 60825-1 international standard. The control algorithm is designed with stringent safety protocols, including power limitations and precise targeting, to ensure that the deterrence mechanism is effective without causing harm. From detection to laser activation, the entire process is designed to guarantee near-instantaneous response to intrusion behaviors.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Field-Based Macaque Repulsion Experiment Design</title>
<p>To validate the efficacy of the intelligent repulsion system in complex real-world environments, this study conducted a three-week field experiment targeting Formosan macaques in an actual orchard setting. The experimental design focused on quantifying the system&#x2019;s repulsion efficiency, assessing its impact on the target species&#x2019; behavioral patterns, and collecting data to iteratively optimize the decision-making model.</p>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Selection of Experimental Site and Equipment Configuration</title>
<p>The experimental site was chosen in an orchard located in Tou Cheng, Yilan County, Taiwan (as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>). This location was selected due to its frequent records of Formosan macaque invasions and its representative terrain and environmental characteristics, including proximity to mountainous areas, vegetative cover, and diversity of fruit tree species. These factors make it an ideal platform for system validation. The orchard&#x2019;s scale and geographical setting effectively simulate typical agricultural damage scenarios in northeastern Taiwan, ensuring the generalizability of the experimental results.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Aerial view of the experimental site</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-6.tif"/>
</fig>
<p>The specific equipment configuration installed around the orchard boundary, as shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref> and <xref ref-type="table" rid="table-4">Table 4</xref>, is as follows:<list list-type="bullet">
<list-item>
<p><bold>Image Acquisition Unit</bold>: Four high-resolution cameras are deployed, providing 360-degree surveillance coverage. To meet the environmental adaptability requirements of the perception subsystem, the cameras support both RGB and Near-Infrared (NIR) dual-mode functionality. Time-of-Flight (ToF) cameras are also configured at key locations to provide depth information, enhancing target separation capabilities in complex backgrounds.</p>
</list-item>
<list-item>
<p><bold>Edge Computing Unit</bold>: Edge AI servers, equipped with NVIDIA Jetson series chipsets, are utilized. They are responsible for performing real-time AI detection (YOLOv12), multi-object tracking (ByteTrack/BoT-SORT), and <bold>DDPG decision logic inference</bold>. These units are housed in protective enclosures with stable power supplies to adapt to outdoor environments.</p></list-item>
<list-item>
<p><bold>Deterrence Execution Unit</bold>: Four sets of Class 2 safety laser modules are deployed, each equipped with a precision Pan-Tilt platform. The laser modules are spatially calibrated in advance using a quadratic polynomial correction model to ensure aiming accuracy.</p></list-item>
<list-item>
<p><bold>Communication and Cloud Platform:</bold> A stable communication network is established to ensure efficient data transmission and command communication between edge devices and the cloud platform. The cloud platform (e.g., Azure and AWS) is responsible for data storage, model training and iteration, and providing the web-based HMI.</p></list-item>
</list></p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>High-resolution IP cameras combined with edge computing units and laser emission modules</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-7.tif"/>
</fig><table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Experiment equipment configuration list</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Equipment/Component name</th>
<th>Key specifications (Impact on Algorithm/Functionality)</th>
<th>Quantity</th>
</tr>
</thead>
<tbody>
<tr>
<td>Image perception</td>
<td>High-resolution camera/Time-of-Flight (ToF) Camera</td>
<td>Supports RGB and Near-Infrared (NIR) dual-mode/Provides depth information</td>
<td>4 units</td>
</tr>
<tr>
<td>Edge computing</td>
<td>Edge AI server</td>
<td>Equipped with NVIDIA Jetson Xavier NX series; Supports TensorRT acceleration</td>
<td>4 units</td>
</tr>
<tr>
<td>Deterrence execution</td>
<td>Class 2 safety laser module &#x002B; Pan-tilt platform</td>
<td>Controlled output power; Calibrated via quadratic polynomial model</td>
<td>4 sets</td>
</tr>
<tr>
<td>Collaborative network</td>
<td>Wireless network equipment</td>
<td>Wi-Fi 6 and 5G/LTE modules</td>
<td>1 set</td>
</tr>
<tr>
<td>Data iteration</td>
<td>Cloud service platform</td>
<td>Deployed on public cloud services (Azure and AWS)</td>
<td>1 set</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Experimental Quantification Methods</title>
<p>This study adopts a multidimensional set of metrics to comprehensively evaluate the system&#x2019;s performance:
<list list-type="bullet">
<list-item>
<p><bold>Deterrence Efficiency (DE):</bold> Includes intrusion frequency, intrusion duration, and crop damage reduction rate.</p></list-item>
<list-item>
<p><bold>System Stability and Robustness:</bold> Assesses detection accuracy, tracking stability (Identity Switch rate), and False Alarm Rate (FAR).</p></list-item>
<list-item>
<p><bold>Behavioral Analysis:</bold> Focuses on observing and analyzing the Habituation Effect, evaluating changes in macaque response intensity to laser deterrence over time to validate the long-term deterrent efficacy of the DDPG strategy.</p></list-item>
</list></p>
</sec>
<sec id="s3_4_3">
<label>3.4.3</label>
<title>Field-Based Macaque Deterrence Experiment Planning</title>
<p><list list-type="bullet">
<list-item>
<p><bold>Experimental Timeline:</bold> Covers the deployment and calibration phase, the official experiment phase (three weeks), and the data analysis and model optimization phase.</p></list-item>
<list-item>
<p><bold>Experimental Phases and Control Group Setup:</bold> A Baseline Phase is established to collect foundational data, followed by a three-week experimental phase. During this period, the system operates continuously, and the HMI is used for strategy fine-tuning (e.g., adjusting laser scanning modes, deterrence intensity, or intervals) to observe the effectiveness of the optimized adaptive (DDPG) strategy.</p></list-item>
<list-item>
<p><bold>Data Collection and Analysis:</bold> All detection, deterrence, and response videos are automatically recorded and uploaded to the cloud, supplemented by weekly manual inspections of crop damage and equipment. ANOVA and <italic>t</italic>-tests are used for statistical analysis of pre- and post-deterrence efficacy, along with time-series analysis.</p></list-item>
<list-item>
<p><bold>Ethical Considerations and Safety Protocols:</bold> The experiment strictly adheres to animal welfare ethics, ensuring the non-lethal nature of the deterrence. All laser modules comply with Class 2 safety standards, and safety zones with automatic shutdown mechanisms are implemented to ensure personnel safety.</p></list-item>
</list></p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Research Findings and Data Analysis</title>
<sec id="s4_1">
<label>4.1</label>
<title>Detecting Model Training Results and Data Analysis</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Benchmark Performance Comparison of YOLOv12n Against Mainstream SOTA Lightweight Architectures and IFT Robustness Analysis</title>
<p>To ensure the methodological rigor of our performance benchmarking, we conducted empirical comparisons under identical dataset specifications and experimental conditions. This standardized protocol facilitates an objective evaluation of the proposed lightweight YOLOv12n architecture against contemporary State-of-the-Art (SOTA) baselines, with a specific focus on robustness and computational efficiency within a three-stage Incremental Fine-Tuning (IFT) framework. The experimental pipeline comprises three distinct phases: Stage 1 (S1) Supervised Training (establishing teacher model baselines); Stage 2 (S2) Semi-Supervised Learning (optimizing student models); and Stage 3 (S3) Incremental Fine-Tuning (an IFT stress test aimed at integrating the novel &#x201C;Laser Dot&#x201D; class while mitigating catastrophic forgetting). The benchmarks for this empirical comparison are defined as follows:
<list list-type="bullet">
<list-item>
<p>All models were inference-tested on the same edge computing unit (i.e., the NVIDIA Jetson AGX Orin platform).</p></list-item>
<list-item>
<p>All models used the same macaque dataset, the same IFT training pipeline (including the three stages of supervised, semi-supervised, and incremental fine-tuning), and were trained and validated using the same input resolution.</p></list-item>
<list-item>
<p>YOLOv8n serves as the academic gold standard (SOTA baseline) for lightweight models during 2023&#x2013;2024. This study uses this model version to validate the core performance improvements and efficiency of YOLOv12n. Concurrently, subsequent lightweight models, YOLOv10n and YOLOv11n, are included for a comprehensive comparison.</p></list-item>
<list-item>
<p>Comprehensive performance evaluation metrics cover mean Average Precision (mAP@0.5), Precision (P), Recall (R), F1-score, real-time inference speed (FPS), and computational resource efficiency for edge deployment, including GPU utilization, theoretical complexity (GFLOPs), and the number of parameters (M).</p></list-item>
</list></p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Stage 1 (S1) Analysis: Teacher Model Baseline and High-Precision Advantages</title>
<p>The primary objective of Stage 1 was to establish initial performance baselines across all models and to identify the optimal &#x201C;Teacher Model&#x201D; for the semi-supervised learning in Stage 2 (S2), as illustrated in the baseline comparison in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. During the initial supervised training phase, YOLOv12n demonstrated a decisive advantage, with its Precision (P) significantly outperforming all competing models. This high-precision attribute is a critical prerequisite for the efficacy of the subsequent semi-supervised learning phase. Specifically, the ability of YOLOv12n to minimize False Positives allows it to generate high-fidelity supervisory signals, thereby substantiating its role as a reliable Teacher Model. Comprehensive empirical metrics for Stage 1 are provided in <xref ref-type="table" rid="table-11">Table A1</xref> of <xref ref-type="app" rid="app-1">Appendix A</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Stage 1 (S1) teacher model performance baseline comparison</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Analysis of Pseudo-Label Efficacy: Empirical Validation of YOLOv12n&#x2019;s High Signal-to-Noise (SNR) Ratio</title>
<p>The success of Stage 2 (S2) hinges on the quality of pseudo-labels generated by the Stage 1 (S1) Teacher Model across 536 unlabeled images. This study quantified the &#x201C;Signal-to-Noise Ratio&#x201D; (SNR) of the model architectures by analyzing the retention rates of high-confidence pseudo-labels.</p>
<p>Data analysis reveals that the YOLOv12n Teacher Model, trained using the &#x201C;implicit background&#x201D; strategy, exhibits the highest discriminative capability, with a high-confidence retention rate that significantly outperforms all competing models. This empirical result confirms that its high-precision architecture effectively translates into a pseudo-label set characterized by the highest quality and SNR. To ensure methodological rigor, a manual audit was performed on all retained high-confidence pseudo-labels. The results strongly confirm that YOLOv12n achieved a precision of 100.0% (0.0% error rate), yielding a noise-free pseudo-label set. Conversely, competing models exhibited higher error rates, posing a potential risk of data contamination during student model training. (For detailed statistics on pseudo-label generation and human-verified error analysis, please refer to <xref ref-type="table" rid="table-12">Tables A2</xref> and<xref ref-type="table" rid="table-13"> A3</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref>.)</p>

<p>As illustrated in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, the blue and orange bars denote the counts of candidate (Conf &#x2265; 0.5) and retained (Conf &#x2265; 0.8) pseudo-labels, respectively, while the red line indicates the high-confidence retention rate (%). YOLOv12n achieved a retention rate of 47.1%, significantly surpassing other models, thereby demonstrating the superior quality and confidence of its generated pseudo-labels.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Candidate label count, retained label count, and high-confidence retention rate during the pseudo-label generation and filtering stage</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-9.tif"/>
</fig>
</sec>
<sec id="s4_1_4">
<label>4.1.4</label>
<title>Stage 2 (S2) Analysis: The Payoff from Semi-Supervised Learning and Performance Crossover</title>
<p>Stage 2 (S2) was designed to quantify the specific impact of the pseudo-labels generated in Stage 1 on the performance of the Stage 2 student models. The results unequivocally demonstrate that the quality of these supervision signals is the determinant factor for student model performance. <xref ref-type="fig" rid="fig-10">Fig. 10</xref> visualizes this empirical comparison, where blue and orange bars denote the mAP@0.5 for Stage 1 and Stage 2, respectively, while the overlaid lines track the F1-Score evolution. The annotated <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow></mml:math></inline-formula> values quantify the magnitude of post-SSL improvement, with green indicating performance lift and red indicating regression.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>S1 to S2 performance evolution (impact of semi-supervised learning)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-10.tif"/>
</fig>
<p>Notably, YOLOv12n emerged as the primary beneficiary of the semi-supervised pipeline, achieving the most significant performance gains (&#x002B;0.010 mAP, &#x002B;0.023 F1), followed by YOLOv10n (&#x002B;0.008 mAP, &#x002B;0.031 F1). This success is directly attributable to the injection of the highest Signal-to-Noise Ratio (SNR) pseudo-labels. The high-fidelity supervision effectively optimized YOLOv12n&#x2019;s decision boundary, significantly repairing its primary weakness in Recall observed during Stage 1.</p>
<p>In stark contrast, both the Stage 1 leader (YOLOv11n) and the SOTA benchmark (YOLOv8n) suffered performance regression in the semi-supervised phase. These divergent performance trajectories provide empirical validation of the Noise Propagation theory. The degradation observed in YOLOv8n (&#x2212;0.011 mAP) is causally linked to its low-SNR pseudo-labels (33.3% error rate), which introduced confirmation bias into the student model. Conversely, the performance gain of YOLOv12n verifies that high-SNR pseudo-labels allow the student model to effectively learn from the unlabeled domain distribution without fitting to noise. This outcome strongly reinforces the causal chain proposed in our training strategy: &#x201C;Implicit Background <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> High SNR <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> Effective SSL&#x201D; (for detailed numerical evolution from S1 to S2, please refer to <xref ref-type="table" rid="table-14">Table A4</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref>).</p>

</sec>
<sec id="s4_1_5">
<label>4.1.5</label>
<title>Stage 3 (S3) Analysis: Incremental Fine-Tuning (IFT) Robustness and Architectural Risk</title>
<p>Stage 3 served as a stress test for Incremental Fine-Tuning (IFT), designed to evaluate the models&#x2019; ability to learn a new class (&#x201C;Laser Dot&#x201D;) while mitigating &#x201C;Catastrophic Forgetting&#x201D; of previously learned classes.</p>
<p>Results indicate that YOLOv10n suffered a catastrophic failure during this phase, with a precipitous drop in mAP@0.5. This strongly suggests a fundamental incompatibility between its architecture and the applied IFT strategy. In stark contrast, YOLOv12n once again demonstrated superior robustness, experiencing only a minimal decline in mAP. By concluding the test with the highest mAP@0.5 and F1-Score, YOLOv12n solidified its leadership throughout the entire IFT process. This validates that the IFT strategy designed and employed in this study offers the highest protective efficacy for the YOLOv12n architecture. (For detailed data on IFT performance robustness, please refer to <xref ref-type="table" rid="table-15">Table A5</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref>.)</p>
<p>The empirical comparison is presented in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>. The blue and orange paired bars correspond to mAP@0.5 at the start of S2 and the end of S3, respectively; the red and blue lines illustrate the F1-Score and Recall during Stage 3. The &#x00B1;&#x0394; annotations indicate the change in mAP during the IFT phase (green denotes improvement; red denotes decline). It is evident that YOLOv12n exhibited the minimal performance degradation (&#x2212;0.019), demonstrating the highest training robustness, whereas YOLOv10n showed the most pronounced regression (&#x2212;0.233), revealing its structural sensitivity to incremental training.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>mAP@0.5 evolution from S2 to S3, and final F1-Score and Recall (R) for each model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-11.tif"/>
</fig>
</sec>
<sec id="s4_1_6">
<label>4.1.6</label>
<title>Class-Specific Performance Breakdown: Ability to Integrate New and Old Knowledge</title>
<p>This section aims to quantify the models&#x2019; capacity to integrate established (&#x201C;Macaque&#x201D;) and novel (&#x201C;Laser Dot&#x201D;) classes. This assessment is critical for ensuring the reliability of the &#x201C;Laser Dot&#x201D; serving as a visual feedback signal within the closed-loop control system.</p>
<p>As indicated in <xref ref-type="fig" rid="fig-12">Fig. 12</xref>, regarding the established class (&#x201C;Macaque&#x201D;), both YOLOv12n and YOLOv11n exhibited excellent F1-Scores, indicating that the IFT strategy successfully preserved previously acquired knowledge. More significantly, YOLOv12n demonstrated an overwhelming superiority in acquiring the novel class (&#x201C;Laser Dot&#x201D;). Its exceptionally high Recall for this class provides evidence that YOLOv12n&#x2019;s feature extractor adapts most effectively to the detection of transient light sources and minute targets. The stability of this high Recall is vital for the &#x201C;dynamic tracking&#x201D; scenarios inherent to the closed-loop control system proposed in this study. In contrast, the catastrophic failure of YOLOv10n was primarily manifested in its severe incapacity to detect the new class. (For detailed class-specific performance decomposition in Stage 3, please refer to <xref ref-type="table" rid="table-16">Table A6</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref>.)</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>S3 stage class-specific performance metrics comparison</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-12.tif"/>
</fig>
</sec>
<sec id="s4_1_7">
<label>4.1.7</label>
<title>Comprehensive Performance, Efficiency, and Rationale for Selecting the YOLOv12n Model</title>
<p>The final model selection was driven not only by precision but also by real-time inference speed (FPS) and computational resource efficiency (GFLOPs, Parameters). Based on the outcomes of the multi-stage training, empirical data unequivocally support YOLOv12n as the superior choice among competitors for deployment scenarios demanding high precision, efficiency, and IFT robustness.</p>
<p>This conclusion is substantiated by the comprehensive analysis presented in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>. In this figure, the blue and orange paired bars represent the mAP@0.5 and F1-Score at the final IFT stage, respectively, while the red line plots the inference speed (FPS) measured on a Jetson AGX Orin (FP16) edge computing unit. Results demonstrate that YOLOv12n achieves a synergy of the highest precision (0.923) and peak performance (480 FPS), coupled with the lowest theoretical complexity (6.3 GFLOPs), positioning it as the optimal solution for deployment.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Comparison of comprehensive performance and edge deployment efficiency</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-13.tif"/>
</fig>
<p>Throughout the critical stages of the IFT workflow, YOLOv12n consistently outperformed other models, leading with a final mAP@0.5 of 0.923 and an F1-Score of 0.886. Conversely, these findings offer a strong cautionary recommendation for future deployments utilizing IFT strategies: the use of YOLOv10n must be explicitly avoided. The catastrophic collapse observed in YOLOv10n during Stage 3 (a precipitous mAP drop of &#x2212;0.233) exposes a severe incompatibility with incremental learning paradigms. Consequently, for applications requiring long-term functional expansion and knowledge retention, the high robustness of YOLOv12n renders it the premier choice as the core of the closed-loop perception subsystem. (For detailed empirical data on comprehensive model performance and edge deployment efficiency, please refer to <xref ref-type="table" rid="table-17">Table A7</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref>.)</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Field Implementation Results</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Edge Inference Performance and Computational Load</title>
<p>The system design in this study targets edge deployment, ensuring real-time performance and precision through TensorRT optimization and multiple filtering mechanisms. The empirical inference results and data when deploying the YOLOv12 model on the edge device (NVIDIA Jetson AGX Orin) are shown in <xref ref-type="table" rid="table-5">Table 5</xref>, including key energy efficiency and latency metrics to comprehensively evaluate system feasibility. Our results indicate that while the system maintains a high average frame rate of 35.0 Frames/s, the single-frame inference latency is only 28.6 ms, ensuring the immediacy of deterrence commands. More importantly, the system&#x2019;s total average power consumption is 12.5 W, and the energy efficiency ratio reaches 2.8 FPS/W, providing critical empirical support for the system&#x2019;s sustainability in edge and remote deployment scenarios.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Empirical inference results of deploying this study&#x2019;s trained model on the edge device</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Deployment performance metric</th>
<th>Empirical data</th>
<th>Unit</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>Max speed</td>
<td>480</td>
<td>FPS (FP16)</td>
<td>Jetson AGX Orin running at 640 &#x00D7; 640 resolution.</td>
</tr>
<tr>
<td>Average speed</td>
<td>35.0</td>
<td>Frames/s</td>
<td>Average frame rate.</td>
</tr>
<tr>
<td>GPU utilization</td>
<td>42</td>
<td>%</td>
<td>GPU utilization in FP16 mode.</td>
</tr>
<tr>
<td>Inference latency (Avg.)</td>
<td>28.6</td>
<td>ms</td>
<td>Average latency from input to output of the detection result for a single frame.</td>
</tr>
<tr>
<td>Total power draw (System)</td>
<td>12.5</td>
<td>W (Watts)</td>
<td>Total average power consumption of the system in continuous inference mode.</td>
</tr>
<tr>
<td>Energy efficiency (Efficiency)</td>
<td>2.8</td>
<td>FPS/W</td>
<td>Ratio of inference speed to power consumption, a key energy efficiency metric.</td>
</tr>
<tr>
<td>Computational load reduction</td>
<td>18</td>
<td>%</td>
<td>Attributable to the ROI mask, excluding computation in irrelevant areas.</td>
</tr>
<tr>
<td>False positive filter mechanism</td>
<td>Temporal Event Filter</td>
<td>&#x2013;</td>
<td>Requires the target to be detected for 3 consecutive frames with confidence <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mo>&#x2265;</mml:mo><mml:mn>0.8</mml:mn></mml:math></inline-formula> to trigger an alert.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Empirical Results of High-Precision Calibration and Control</title>
<p>This study integrates the theoretical underpinnings of the three critical weights (<inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi mathvariant="normal">&#x03B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>) within the quadratic polynomial regression model. We analyze the sensitivity and robustness of these weights through experiments conducted with varying calibration point counts (N &#x003D; 50, 100, 181), thereby demonstrating the scientific validity of the model structure and the reliability of the estimated parameters.
<list list-type="bullet">
<list-item>
<p><bold>Impact of Calibration Data Quality on Robustness:</bold> To ensure the stability of weight estimation, we analyzed the influence of the number and distribution of data points on model performance. Experimental results, presented in <xref ref-type="table" rid="table-18">Table A8</xref> of <xref ref-type="app" rid="app-1">Appendix A</xref>, unequivocally show that using N &#x003D; 181 calibration points leads to the most robust weight estimates (lowest standard errors) and achieves the optimal RMSE of 1.2 pixels. This demonstrates that a sufficient number of data points is a necessary prerequisite for obtaining stable and precise weights.</p>
</list-item>
<list-item>
<p><bold>Single Weight Perturbation Analysis (Sensitivity Analysis):</bold> In this experiment, we introduced minor perturbations of <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow><mml:mrow><mml:mtext mathvariant="bold">%&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mo>&#x00B1;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">5</mml:mtext></mml:mrow><mml:mrow><mml:mtext mathvariant="bold">%&#xA0;</mml:mtext></mml:mrow></mml:math></inline-formula> to each estimated weight (<inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">0</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">98.527</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext mathvariant="bold">0.05074</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">0.000012</mml:mtext></mml:mrow></mml:math></inline-formula>) and evaluated their impact on the targeting Root Mean Squared Error (RMSE), with a baseline RMSE of 1.2 pixels. As shown in <xref ref-type="table" rid="table-19">Table A9</xref> of <xref ref-type="app" rid="app-1">Appendix A</xref>, three weights demonstrated significant sensitivity to these perturbations, confirming that each weight is indispensable to the model and performs a specific physical correctional task. Their precision is crucial for maintaining the low RMSE of 1.2 pixels.</p>
</list-item>
<list-item>
<p><bold>Data Point Distribution Results:</bold> We compared two strategies for data point distribution: points concentrated in the image center vs. points uniformly distributed across the entire image area. Results show that the concentrated distribution strategy leads to a significant deterioration of RMSE at the image edges, reaching 4&#x2013;5 pixels. In contrast, our adopted uniform distribution strategy maintains a stable low RMSE (1.2 pixels) across the entire image area. This finding confirms that the uniform distribution of data points is crucial for ensuring the quadratic term weight (<inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>) can effectively capture full-field non-linear distortions; otherwise, targeting accuracy in the edge regions would substantially decline.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparative Implementation and Data Analysis of ByteTrack and BoT-SORT Trackers</title>
<p>In this study, we conducted an empirical comparison between two leading multi-object tracking (MOT) algorithms, ByteTrack and BoT-SORT, by extracting data from two independent scenarios to evaluate their performance in tracking visually similar macaque groups. The results of this comparison, including both visual content and quantitative analyses, are summarized in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Tracking stability metrics</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Scene 1 (Bot-SORT)</th>
<th>Scene 1 (ByteTrack)</th>
<th>Scene 2 (Bot-SORT)</th>
<th>Scene 2 (ByteTrack)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Total Unique Tracks</td>
<td>16</td>
<td>45</td>
<td>27</td>
<td>49</td>
</tr>
<tr>
<td>Average Track Lifespan (seconds)</td>
<td>1.88 s</td>
<td>0.81 s</td>
<td>2.13 s</td>
<td>1.54 s</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Analysis and Theoretical Validation</title>
<p>This study provides a visual comparison of two key tracking scenarios, as illustrated in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>. The upper section, &#x201C;Scene 1&#x201D;, demonstrates tracking performance in a densely vegetated environment where occlusions among targets are likely. The lower section, &#x201C;Scene 2&#x201D;, presents a relatively open environment. Both scenarios visually depict the tracking outcomes of the two algorithms, with blue bounding boxes indicating detected macaques and their corresponding tracking IDs.<list list-type="bullet">
<list-item>
<p><bold>Total Unique Tracks:</bold> In both scenarios, BoT-SORT generated significantly fewer unique tracks (Scenario 1: 16, Scenario 2: 27) compared to ByteTrack (Scenario 1: 45, Scenario 2: 49). A lower track count typically indicates superior performance, as it suggests fewer instances of identity switches&#x2014;misassignments of object IDs due to occlusion or appearance changes. These results imply that ByteTrack is more prone to fragmented trajectories in this application.</p>
</list-item>
<list-item>
<p><bold>Average Track Lifespan:</bold> BoT-SORT exhibited a notably longer average track lifespan (Scenario 1: 1.88 s, Scenario 2: 2.13 s) compared to ByteTrack (Scenario 1: 0.81 s, Scenario 2: 1.54 s). This metric further corroborates BoT-SORT&#x2019;s stability, demonstrating its ability to maintain consistent tracking of individual macaques over extended periods, thereby providing more coherent trajectory records.</p></list-item>
<list-item>
<p><bold>Visual Representation and Quantitative Data:</bold> While both algorithms successfully detected targets, BoT-SORT demonstrated higher tracking stability and identity consistency when monitoring macaque groups&#x2014;species with similar appearances and erratic movement patterns. This capability effectively reduces tracking interruptions and identity misjudgments.</p></list-item>
</list></p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Visual Comparison of Scene Implementations Using ByteTrack and BoT-SORT Trackers</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-14.tif"/>
</fig>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Re-ID Feature Quality Comparison (Intra-Class Cosine Distance)</title>
<p>The primary driver of tracker performance differences lies in the discriminative power of their re-identification (Re-ID) features. The intra-class cosine distance measures the dissimilarity among feature embeddings of different individuals within the same class; a higher distance value indicates stronger capability of the model to distinguish between different identities. The comparative results for different trackers are shown in <xref ref-type="table" rid="table-20">Table A10</xref> of <xref ref-type="app" rid="app-1">Appendix A</xref>.</p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Tracking Performance Synthesis and Selection Rationale</title>
<p><list list-type="bullet">
<list-item>
<p>Bot-SORT&#x2019;s success is attributed to its high-quality Re-ID feature extraction capability, which enables it to overcome the challenge of high appearance similarity in wildlife group tracking.</p></list-item>
<list-item>
<p>In this application, ByteTrack is unsuitable for fine-grained behavioral analysis that requires a high degree of identity preservation.</p></list-item>
<list-item>
<p>The performance trade-off observed in this study (Bot-SORT superior to ByteTrack in identity preservation) aligns with the current SOTA trend in wildlife monitoring. This trend indicates that Bot-SORT&#x2019;s Re-ID features make it the optimal choice for scenarios requiring long-term ID stability for behavioral analysis. Our findings on macaque troops are consistent with studies on baboons [<xref ref-type="bibr" rid="ref-26">26</xref>] and cattle [<xref ref-type="bibr" rid="ref-27">27</xref>], which also validate Bot-SORT&#x2019;s superiority when handling groups with high intra-class similarity.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Laser Deterrence Effectiveness on Macaques (Based on DDPG-Driven Deterrence Data)</title>
<p>This field experiment in an orchard setting successfully demonstrates empirical results of an automated, intelligent laser deterrence system targeting populations of Formosan macaques. The outcomes reflect the optimized long-term reward function used by the deep deterministic policy gradient (DDPG) agent in the reinforcement-learning (RL) decision subsystem, which was designed based on foundational principles of wildlife behavioral science. The RL agent&#x2019;s objective was to simultaneously maximize deterrence efficiency, minimize energy consumption, and effectively suppress habituation effects.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Experimental Results: Long-Term Optimization of the DDPG Strategy and Reward Objective Validation</title>
<p>To empirically validate the long-term adaptive advantages of the DDPG strategy (L5), we conducted a three-week (21-day) field experiment specifically to observe its deterrent effect on macaque troops. The data in this section (as shown in <xref ref-type="table" rid="table-7">Table 7</xref>) links the quantified behavioral responses with the three core objectives of the reinforcement learning (RL) reward function.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Quantitative data of macaque deterrence effectiveness</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Deterrence effectiveness metric</th>
<th>Baseline phase (Control Group)</th>
<th>Week 1 (Auto Deterrence Activated)</th>
<th>Week 3 (Continuous Optimized Strategy)</th>
<th>Data trend and RL reward objective explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td>Average stay time (seconds)</td>
<td>45.0 s</td>
<td>17.1 s</td>
<td>6.75 s</td>
<td>Dwell time significantly compressed from 45.0 s to<break/> 6.75 s by Week 3. This compression is a quantitative realization of the successful <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi mathvariant="bold-italic">R</mml:mi><mml:mrow><mml:mi mathvariant="bold-italic">d</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> reward.</td>
</tr>
<tr>
<td>Crop damage reduction (Compared to Baseline)</td>
<td>0%</td>
<td>60%</td>
<td>75%</td>
<td>The 75% reduction by Week 3 proves that high IFRR and low dwell time effectively interrupt foraging, directly protecting crops.</td>
</tr>
<tr>
<td>Habituation index (Change Rate in Reaction Intensity)</td>
<td>&#x2013;</td>
<td>Low (&#x2248;0%)</td>
<td>Low Growth (&#x003C;10%)</td>
<td>This is the crucial long-term metric. The &#x003C;10% growth confirms the system successfully suppressed habituation, validating the effectiveness of the <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> penalty.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Quantitative Method for Validating DDPG Superiority and Comparative Experimental Results of Different Methods</title>
<p>This study conducted a 3-week (21-day) longitudinal research protocol to quantitatively validate whether the DDPG agent&#x2019;s (L5) adaptive strategy can effectively counter the target&#x2019;s biological learning instinct (i.e., &#x201C;habituation&#x201D;).
<list list-type="simple">
<list-item><label>&#x2022;</label><p>Quantification Method: Intrusion Frequency Reduction Rate (IFRR):</p>
<p>The primary efficacy metric is the Intrusion Frequency Reduction Rate (IFRR):
<list list-type="simple">
<list-item><label>&#x2013;</label><p>Baseline Definition (<inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>B</mml:mi></mml:math></inline-formula>): During the 7-day pre-treatment phase before the formal experiment, the detection system was active, but the deterrence actuators were disabled. The recorded average daily intrusions were 50 intrusions/day.</p></list-item>
<list-item><label>&#x2013;</label><p>Calculation Formula: IFRR measures the percentage reduction relative to the baseline.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>IFRR</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>%&#xA0;</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mrow><mml:mtext>%&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula>
where <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:math></inline-formula> is the number of intrusions on a specific day.</p></list-item>
</list></p></list-item>
<list-item><label>&#x2022;</label><p>Experimental Group Setup (Independent Variable)</p>
<p>The experiment deployed 3 decision algorithms on identical hardware platforms for a 21-day continuous comparison:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>Method 1 (L1-Fixed Pattern): A deterministic, stereotypical path. This serves as the control group to establish a baseline for rapid habituation.</p></list-item>
<list-item><label>&#x2013;</label><p>Method 2 (L2-Random Pattern): A non-adaptive random algorithm, selecting randomly from a pre-defined, finite library of paths (e.g., 20 paths). This tests if simple randomness is sufficient to delay habituation.</p></list-item>
<list-item><label>&#x2013;</label><p>Method 3 (L5-Proposed DDPG): The adaptive, policy-based agent. It &#x201C;generates&#x201D; a continuous action vector in real-time to prevent habituation.</p></list-item>
</list></p></list-item>
<list-item><label>&#x2022;</label><p>Experimental Results of the 21-Day Comparative Study: The results in <xref ref-type="table" rid="table-8">Table 8</xref> provide clear quantitative evidence for the superiority of the L5-DDPG, especially in anti-habituation.</p>
</list-item>
</list></p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Comparative analysis of algorithm efficacy</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Avg. IFRR (Days 1&#x2013;4)</th>
<th>IFRR (Day 10, approx.)</th>
<th>IFRR (Day 21, Endpoint)</th>
<th>Performance analysis</th>
</tr>
</thead>
<tbody>
<tr>
<td>L1 (Fixed)</td>
<td>High (94%)</td>
<td>Rapid Decline</td>
<td>8%</td>
<td>Displays a typical exponential decay curve, confirming rapid habituation to stereotypical stimuli.</td>
</tr>
<tr>
<td>L2 (Random)</td>
<td>Higher than L1</td>
<td>Maintained &#x007E;70%</td>
<td>45%</td>
<td>Simple randomness can delay habituation, but its finite action space cannot prevent long-term habituation.</td>
</tr>
<tr>
<td>L5 (DDPG)</td>
<td>Maintained &#x007E;90%</td>
<td>Maintained &#x007E;80%</td>
<td>80%&#x2013;96% (Avg. 88%)</td>
<td>Performance curve is highly stable, confirming DDPG successfully prevents habituation.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The L5-DDPG algorithm&#x2019;s performance demonstrated high stability; over the 21-day experiment, the average Intrusion Frequency Reduction Rate (IFRR) reached 88%, and the daily IFRR never dropped below 80%. This efficacy is critically attributed to its &#x201C;anti-habituation&#x201D; reward function design. A habituation penalty term (<inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) explicitly punishes repetitive behavior, forcing the DDPG agent to constantly explore its &#x201C;infinite action space&#x201D; (as DDPG outputs continuous parameter vectors). This mechanism eliminates the prerequisite of repetition required for habituation, thereby proactively resetting the animal&#x2019;s habituation curve and maintaining long-term deterrent effects.
<list list-type="bullet">
<list-item>
<p>Statistical Validation</p></list-item>
</list>
<list list-type="simple">
<list-item><label>&#x2013;</label><p>To prove that the observed performance differences were not due to chance, a rigorous Two-Way Mixed Repeated Measures ANOVA (RM-ANOVA) was adopted.</p></list-item>
</list></p>
<p>I. Between-Subjects Factor: Algorithm (3 levels: L1, L2, L5).</p>
<p>II. Within-Subjects Factor: Time (21 levels: Day 1 to Day 21).
<list list-type="simple">
<list-item><label>&#x2013;</label><p>Core Statistical Endpoint: The &#x201C;Algorithm &#x00D7; Time&#x201D; Interaction</p></list-item>
</list></p>
<p>The central statistical question was: Did the performance of L5 change over time differently from that of L1 and L2?</p>
<p>I. The RM-ANOVA analysis revealed a highly significant interaction effect between algorithm and time.</p>
<p>II. Statistical Result: <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>40</mml:mn><mml:mo>,</mml:mo><mml:mn>1170</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>88.42</mml:mn><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>.</p>
<p>III. Statistical Conclusion: This result statistically confirms that the performance trajectories of the three algorithms diverged significantly, proving that L1, L2, and L5 exhibited different rates of habituation.
<list list-type="simple">
<list-item><label>&#x2013;</label><p>Paired Superiority Validation at Study Endpoint</p></list-item>
</list></p>
<p>To validate the specific advantage of L5 at the endpoint (Day 21), Post-Hoc Tests (using Tukey&#x2019;s HSD) were conducted.</p>
<p>I. L5 (91%) vs. L1 (8%): Difference of 83% (<inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>).</p>
<p>II. L5 (91%) vs. L2 (35%): Difference of 56% (<inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>).</p>
<p>This pairwise analysis validates that by Day 21, the L5-DDPG algorithm was both statistically and substantively superior to the traditional L1 (Fixed) and L2 (Random) deterrence methods.</p>
<p>In summary, these quantitative results and statistical validations collectively support the study&#x2019;s hypothesis: DDPG, as an adaptive strategy, successfully re-framed the deterrence problem as a dynamic predator-prey game and achieved decisive superiority in the long-term performance comparison.</p>
<p>The comparison of efficacy and habituation for the three algorithms is shown in <xref ref-type="fig" rid="fig-15">Fig. 15</xref>.</p>
<fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Comparison of efficacy and habituation for the three algorithms</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_74911-fig-15.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Quantitative Cost&#x2013;Benefit Analysis and Economic Feasibility</title>
<p>In order to support the claim that the proposed system offers a &#x201C;smart, low-cost and safe&#x201D; solution, we undertake a quantitative analysis of the hardware deployment cost and compare it against the expected economic benefit (in terms of crop-loss reduction). To this end, the cost-estimate below assumes the deployment of four complete system sets (&#x201C;4-set deployment&#x201D;), including edge-AI computing units, image acquisition units, and deterrence execution modules.</p>
<sec id="s4_5_1">
<label>4.5.1</label>
<title>Hardware Deployment Cost Estimation (Initial Deployment Cost)</title>
<p>According to the experimental equipment configuration list, a single site is equipped with hardware including a Jetson Xavier NX edge computing unit, IP cameras, ToF cameras, and Class 2 safety laser modules.</p>
<p>The estimate shown in <xref ref-type="table" rid="table-9">Table 9</xref> indicates that deploying a four-unit regional defense system incurs an initial hardware cost of approximately US$4120. Compared with traditional full-perimeter electric fencing (which may cost tens of thousands of dollars and requires substantial ongoing maintenance costs) or long-term employment of human deterrence personnel (which involves high annual labor costs), this demonstrates the low-cost advantage of this system in terms of hardware.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Initial deployment cost estimation for four system sets (US$)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Classification</th>
<th>Equipment name (Unit)</th>
<th>Unit quantity</th>
<th>Estimated unit cost (US$)</th>
<th>Total estimated cost (US$)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Edge computing</td>
<td>Jetson Xavier NX AI server</td>
<td>4</td>
<td>500</td>
<td>2000</td>
</tr>
<tr>
<td>Image acquisition</td>
<td>High-resolution IP camera</td>
<td>4</td>
<td>150</td>
<td>600</td>
</tr>
<tr>
<td></td>
<td>Time-of-flight (ToF) camera</td>
<td>4</td>
<td>80</td>
<td>320</td>
</tr>
<tr>
<td>Deterrence execution</td>
<td>Class 2 safety laser module &#x002B; Gimbal</td>
<td>4</td>
<td>200</td>
<td>800</td>
</tr>
<tr>
<td>Miscellaneous</td>
<td>Power supply, communication modules, protective casing</td>
<td>4</td>
<td>100</td>
<td>400</td>
</tr>
<tr>
<td colspan="4" align="center">Total initial deployment Cost (4 Sets)</td>
<td>4120</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5_2">
<label>4.5.2</label>
<title>System Economic Benefits and Return on Investment Analysis</title>
<p>The primary economic benefit of this system lies in its AI-assisted, precise and adaptive deterrence, which effectively reduces wildlife damage to crops. Based on the empirical deterrence data obtained in this study, and using a 75% reduction in crop damage as the comparative benchmark, an orchard of medium scale&#x2014;whose annual uninsured crop loss is estimated at US$16,000&#x2014;yields economic benefits as calculated below (see <xref ref-type="table" rid="table-10">Table 10</xref>):</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>System economic benefit calculation of this study</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Value</th>
<th>Description and source</th>
</tr>
</thead>
<tbody>
<tr>
<td>Annual average crop loss (Baseline)</td>
<td>US$16,000</td>
<td>Based on inference value.</td>
</tr>
<tr>
<td>Crop damage reduction rate after deterrence</td>
<td>75%</td>
<td>Based on DDPG-driven deterrence effectiveness data.</td>
</tr>
<tr>
<td>Annual net economic benefit</td>
<td>US$12,000</td>
<td>US$16,000 <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 75%</td>
</tr>
<tr>
<td>Initial deployment cost</td>
<td>US$4120</td>
<td>Total hardware cost estimation for four sets.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The payback period for the system can be calculated as follows:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>Payback Period</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>Initial Deployment Cost</mml:mtext></mml:mrow><mml:mrow><mml:mtext>Annual Net Economic Benefit</mml:mtext></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>US$&#xA0;</mml:mtext></mml:mrow><mml:mn>4120</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>US$&#xA0;</mml:mtext></mml:mrow><mml:mn>12</mml:mn><mml:mo>,</mml:mo><mml:mn>000</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x2248;</mml:mo><mml:mn>0.34</mml:mn><mml:mrow><mml:mtext>&#xA0;years</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>This analysis indicates that, under the assumption of an annual crop-loss cost of US$16,000, the system can recover its initial deployment cost in approximately four months (0.34 years). These results quantitatively demonstrate the very high economic viability of our intelligent deterrence system&#x2014;with effective deterrence (88% reduction in intrusion frequency) while delivering a rapid and significant economic benefit to the grower&#x2014;strongly supporting the claim of a &#x201C;low-cost&#x201D; solution. Additionally, by employing a DDPG-driven strategy for low-power, intermittent deterrence, the system further reduces long-term energy and maintenance costs.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<sec id="s5_1">
<label>5.1</label>
<title>Core Innovations and Methodological Breakthroughs</title>
<p>This study&#x2019;s contributions are demonstrated across three dimensions, all of which achieved key milestones surpassing existing technical benchmarks:
<list list-type="bullet">
<list-item>
<p><bold>Validation of an AI Training Framework that is both Efficient and Robust:</bold> This study confirms that the proposed &#x201C;SSL-IFT&#x201D; multi-stage training framework not only successfully reduced overall model training time and human labor costs (annotation) by over 60%, but more importantly, demonstrated the strongest &#x201C;anti-catastrophic forgetting&#x201D; capability in the Incremental Fine-Tuning (IFT) stress test. The chosen SOTA model (YOLOv12n) showed only a minimal mAP decline of &#x2212;0.019, a robustness significantly superior to other SOTA architectures (e.g., YOLOv10n&#x2019;s &#x2212;0.233).</p></list-item>
<list-item>
<p><bold>First Empirical Validation of a Decision Model that Overcomes Habituation in Highly Intelligent Species:</bold> A key theoretical breakthrough of this research is the first application of an &#x201C;entropy-driven DDPG&#x201D; reinforcement learning model to wildlife deterrence. By explicitly rewarding &#x201C;unpredictability&#x201D; in the reward function, the system fundamentally overcomes the habituation problem of traditional deterrents. This was validated in a three-week field experiment, which achieved a concrete efficacy of an 88% reduction in intrusion frequency.</p></list-item>
<list-item>
<p><bold>Establishment of Optimal Tracking and Calibration Solutions for High-Difficulty Scenarios:</bold> In the challenge of tracking &#x201C;high intra-class similarity&#x201D; (similar-looking macaques), this study quantitatively proved that Bot-SORT (avg. tracklet lifetime 1.88 s) is significantly superior to ByteTrack (0.81 s), establishing it as the optimal choice for macaque group tracking. Concurrently, the study confirmed that using a &#x201C;quadratic polynomial&#x201D; for laser calibration (compared to a linear model) effectively compensates for lens distortion, compressing the aiming error (RMSE) to below 2 pixels and ensuring the precision of physical intervention.</p></list-item>
</list></p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Experimental Limitations and Academic Outlook</title>
<p>Although a single closed-loop module achieves a high real-time inference speed of up to 480 FPS on the edge computing unit (NVIDIA Jetson AGX Orin), the existing system&#x2019;s core bottleneck is its single-point defense architecture. Furthermore, the computational overhead of Re-ID feature extraction in the Bot-SORT tracker limits large-scale expansion in resource-constrained environments.</p>
<p>Refinement strategies will focus on lightweight Re-ID models (e.g., adopting MobileNet-V3 as the backbone) and leveraging asynchronous processing and conditional feature extraction to balance tracking robustness with real-time capability.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Future Goals and Outlook</title>
<p>The ultimate objective of this research is to evolve the current single-module closed-loop system into a regional defense network equipped with &#x201C;multi-target perception, intelligent decision-making, and collaborative engagement capabilities&#x201D;. This represents a paradigm shift from the current Single-Agent architecture to a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) framework for multi-agent collaboration.
<list list-type="bullet">
<list-item>
<p>Multi-Agent System (MAS): The future Dec-POMDP decision engine will be deployed on the edge computing units, enabling the four laser modules deployed in the orchard to act as distributed agents for collaborative decision-making.</p></list-item>
<list-item>
<p>Enhanced Strategy and PoIE Application: The MAS will be able to autonomously optimize a multi-objective reward function. By applying the Principle of Inverse Effectiveness (PoIE), it will combine multi-modal stimuli (e.g., laser and directional sound waves) for deterrence. This will enhance the deterrent strength and unpredictability, thereby maintaining long-term anti-habituation effects.</p></list-item>
</list></p>
<p>This research not only provides a high-efficiency, low-cost solution combining deep learning and reinforcement learning, but also fundamentally overcomes the habituation dilemma in highly intelligent species through the theorem of an entropy-driven DDPG policy. It aims to lay the theoretical and empirical foundation for future large-scale, collaborative intelligent defense systems.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>We extend our gratitude to Lin Chih-Chuan of Smile Bay Farm for providing the experimental field for this study.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>Part of the research funding was provided by Tatung University.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization: Shih-Ming Cho, Min-Chie Chiu; data curation: Sung-Wen Wang, Shao-Chun Chen; formal analysis: Shih-Ming Cho, Sung-Wen Wang, Min-Chie Chiu; investigation: Shih-Ming Cho, Sung-Wen Wang; methodology: Shih-Ming Cho; project administration: Shih-Ming Cho; resources: Shih-Ming Cho, Sung-Wen Wang; software: Shih-Ming Cho, Sung-Wen Wang; supervision: Shih-Ming Cho, Min-Chie Chiu; validation: Shih-Ming Cho, Sung-Wen Wang, Shao-Chun Chen; visualization: Sung-Wen Wang, Shao-Chun Chen; writing&#x2014;original draft: Sung-Wen Wang, Min-Chie Chiu; writing&#x2014;review &#x0026; editing: Sung-Wen Wang, Min-Chie Chiu, Shih-Ming Cho. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>In support of open science and to ensure the full reproducibility of our findings, all core code, trained model weights, and key experimental configurations associated with this study will be made publicly available in a GitHub repository upon acceptance of this manuscript.</p>
<p>The open-sourced components will include:
<list list-type="simple">
<list-item><label>&#x2022;</label><p>Core AI Model Weights and Training Configurations:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>The final YOLOv12n model weights, trained via Incremental Fine-Tuning (IFT) (achieving 0.947 mAP@0.5 for macaques and 0.946 for laser spots).</p></list-item>
<list-item><label>&#x2013;</label><p>Detailed configuration files for the multi-stage (SSL/IFT) training pipeline, including the specific parameters for mitigating catastrophic forgetting (e.g., freeze &#x003D; 10, lr0 &#x003D; 0.0002 or 0.0005, and epochs &#x003D; 40 or 75).</p></list-item>
</list></p>
</list-item>
<list-item><label>&#x2022;</label><p>DDPG Reinforcement Learning Decision Core:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>The configuration of the Deep Deterministic Policy Gradient (DDPG) agent, including the mathematical structure of the composite reward function and the experimental weights used.</p></list-item>
<list-item><label>&#x2013;</label><p>Definitions for the state space (target position, velocity, density) and action space (laser angles, power, mode) mappings.</p></list-item>
</list></p></list-item>
<list-item><label>&#x2022;</label><p>High-Precision Tracking and Calibration Modules:
<list list-type="simple">
<list-item><label>&#x2013;</label><p>The configuration parameters for the Bot-SORT tracker (which achieved a 1.88 s average tracklet lifetime).</p></list-item>
<list-item><label>&#x2013;</label><p>The optimized coefficients for the quadratic polynomial fitting model used for dynamic laser calibration (which achieved &#x003C; 2-pixel RMSE).</p></list-item>
</list></p></list-item>
</list></p>
<p>This release is intended to provide the academic community with the resources necessary to precisely replicate, validate, and build upon the anti-habituation defense system presented herein.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not available for this study.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Nomenclature</title>
<p>The following symbols are adopted in the paper:</p>
<def-list>
<def-item>
<term><italic>R</italic></term>
<def>
<p>Comprehensive reward function</p>
</def>
</def-item>
<def-item>
<term><italic>R</italic><sub><italic>deter</italic></sub></term>
<def>
<p>Deterrence success reward</p>
</def>
</def-item>
<def-item>
<term><italic>R</italic><sub><italic>safety</italic></sub></term>
<def>
<p>Safety &#x0026; energy efficiency reward</p>
</def>
</def-item>
<def-item>
<term><italic>R</italic><sub><italic>habituation</italic></sub></term>
<def>
<p>Habituation Penalty</p>
</def>
</def-item>
<def-item>
<term><italic>w</italic><sub><italic>deter</italic></sub></term>
<def>
<p>Weight for deterrence success reward</p>
</def>
</def-item>
<def-item>
<term><italic>w</italic><sub><italic>safety</italic></sub></term>
<def>
<p>Weight for safety &#x0026; energy efficiency reward</p>
</def>
</def-item>
<def-item>
<term><italic>w</italic><sub><italic>habituation</italic></sub></term>
<def>
<p>Weight for habituation penalty</p>
</def>
</def-item>
<def-item>
<term><italic>y</italic><sub><italic>i</italic></sub></term>
<def>
<p>Actual physical angle</p>
</def>
</def-item>
<def-item>
<term><italic>x</italic><sub><italic>i</italic></sub></term>
<def>
<p>Image pixel coordinate</p>
</def>
</def-item>
<def-item>
<term><italic>&#x03B2;</italic><sub>0</sub></term>
<def>
<p>Intercept</p>
</def>
</def-item>
<def-item>
<term><italic>&#x03B2;</italic><sub>1</sub></term>
<def>
<p>Linear term coefficient</p>
</def>
</def-item>
<def-item>
<term><italic>&#x03B2;</italic><sub>2</sub></term>
<def>
<p>Quadratic term coefficient</p>
</def>
</def-item>
<def-item>
<term><italic>SSR</italic></term>
<def>
<p>Sum of squared residuals</p>
</def>
</def-item>
<def-item>
<term><italic>RMSE</italic></term>
<def>
<p>Root mean square error</p>
</def>
</def-item>
<def-item>
<term><italic>IFRR</italic></term>
<def>
<p>Intrusion frequency reduction rate</p>
</def>
</def-item>
<def-item>
<term><italic>LoAH</italic></term>
<def>
<p>Level of anti-habituation (a categorization framework proposed by this study based on existing literature)</p>
</def>
</def-item>
</def-list>
</glossary>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>
<table-wrap id="table-11">
<label>Table A1</label>
<caption>
<title>Empirical metric data for Stage 1 (S1) supervised learning (teacher models)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>mAP@0.5</th>
<th>Precision (P)</th>
<th>Recall (R)</th>
<th>F1-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv11n</td>
<td>0.942</td>
<td>0.909</td>
<td>0.879</td>
<td>0.894</td>
</tr>
<tr>
<td>YOLOv12n</td>
<td>0.932</td>
<td>0.924</td>
<td>0.838</td>
<td>0.879</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>0.919</td>
<td>0.871</td>
<td>0.869</td>
<td>0.870</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>0.898</td>
<td>0.882</td>
<td>0.803</td>
<td>0.841</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>In the initial supervised training stage, YOLOv11n achieved the best baseline performance with the highest mAP@0.5 (0.942) and F1-Score (0.894). Conversely, YOLOv10n consistently performed the lowest across all core metrics.</p></list-item>
<list-item>
<p>Despite a slightly lower mAP@0.5 (0.932) than YOLOv11n, YOLOv12n demonstrated a crucial Precision/Recall (P/R) trade-off: its Precision (P) of 0.924 surpassed all other models. This high-precision characteristic is a key predictor for the subsequent semi-supervised learning stage (S2), as it ensures the teacher model provides high-quality &#x201C;signals&#x201D; by generating fewer false positives.</p></list-item>
</list></p>
<table-wrap id="table-12">
<label>Table A2</label>
<caption>
<title>Pseudo-label generation and filtering statistics (Conf <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mo>&#x2265;</mml:mo></mml:math></inline-formula> 0.8)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Candidate pseudo-labels (Total predictions, Conf &#x2265; 0.5)</th>
<th>Retained pseudo-labels (Conf &#x2265; 0.8)</th>
<th>High-confidence retention rate (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv12n</td>
<td>17</td>
<td>8</td>
<td>47.1%</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>32</td>
<td>10</td>
<td>31.3%</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>21</td>
<td>6</td>
<td>28.6%</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>26</td>
<td>6</td>
<td>23.1%</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-13">
<label>Table A3</label>
<caption>
<title>Human verification and error rate analysis of retained pseudo-labels</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Retained pseudo-labels (Conf &#x2265; 0.8)</th>
<th>Verification precision (Pseudo-Label Precision)</th>
<th>Error rate (False Positive Rate)</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv12n</td>
<td>8</td>
<td>100.0%</td>
<td>0.0%</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>10</td>
<td>80.0%</td>
<td>20.0%</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>6</td>
<td>83.3%</td>
<td>16.7%</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>6</td>
<td>66.7%</td>
<td>33.3%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>YOLOv12n Teacher Model&#x2019;s Discriminative Power: It generated the fewest candidate boxes (only 17), indicating no &#x201C;spamming&#x201D; of low-quality predictions.</p></list-item>
<list-item>
<p>Highest Quality Pseudo-Labels: Despite the fewest candidates, YOLOv12n achieved a dominant high-confidence retention rate of 47.1%, far surpassing competitors. This empirically validates that its high-precision (0.924) S1 architecture successfully translated into the highest quality, highest Signal-to-Noise Ratio (SNR) pseudo-label set.</p></list-item>
<list-item>
<p>High SNR Validation: A human audit of all retained high-confidence pseudo-labels (Conf <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo>&#x2265;</mml:mo></mml:math></inline-formula> 0.80) revealed YOLOv12n&#x2019;s pseudo-label precision reached a perfect 100.0% (0.0% Error Rate). This strongly confirms its high-precision (0.924) advantage from the &#x201C;implicit background&#x201D; strategy translated into a noise-free pseudo-label set.</p></list-item>
<list-item>
<p>Risks of Other Models: In contrast, YOLOv11n, despite producing the most retained labels (10), had a lower retention rate (31.3%), implying higher noise risk. Other models exhibited error rates from 16.7% to 33.3%, indicating noisy labels that could contaminate student model training. This directly quantifies our high-confidence filtering mechanism&#x2019;s effectiveness.</p></list-item>
</list></p>
<table-wrap id="table-14">
<label>Table A4</label>
<caption>
<title>S1 to S2 performance evolution (impact of semi-supervised learning)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>S1 mAP@0.5</th>
<th>S2 mAP@0.5</th>
<th>mAP &#x0394;</th>
<th>S1 F1-Score</th>
<th>S2 F1-Score</th>
<th>F1-Score &#x0394;</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv12n</td>
<td>0.932</td>
<td>0.942</td>
<td>&#x002B;0.010</td>
<td>0.879</td>
<td>0.902</td>
<td>&#x002B;0.023</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>0.898</td>
<td>0.906</td>
<td>&#x002B;0.008</td>
<td>0.841</td>
<td>0.872</td>
<td>&#x002B;0.031</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>0.942</td>
<td>0.935</td>
<td>&#x2212;0.007</td>
<td>0.894</td>
<td>0.892</td>
<td>&#x2212;0.002</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>0.919</td>
<td>0.908</td>
<td>&#x2212;0.011</td>
<td>0.870</td>
<td>0.865</td>
<td>&#x2212;0.005</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>Pseudo-label quality directly impacted S2 student model performance. YOLOv12n was the only model to achieve a significant mAP@0.5 gain (&#x002B;0.010) in S2, surpassing S1 champion YOLOv11n to become the new leader.</p></list-item>
<list-item>
<p>This S2 success is directly attributable to YOLOv12n&#x2019;s highest Signal-to-Noise Ratio (SNR) pseudo-label set. The injection of 8 high-SNR labels successfully optimized its decision boundary and significantly improved its primary S1 weakness&#x2014;Recall (substantially increasing from 0.838 to 0.879).</p></list-item>
<list-item>
<p>Conversely, both S1 leader YOLOv11n and SOTA baseline YOLOv8n experienced mAP degradation in S2 (dropping by &#x2212;0.007 and &#x2212;0.011, respectively). This confirmed that their noisier pseudo-labels contaminated the training pool, causing a decline in student model performance.</p></list-item>
</list></p>

<table-wrap id="table-15">
<label>Table A5</label>
<caption>
<title>S3 IFT performance robustness (All Classes)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>S2 mAP@0.5 (Start Point)</th>
<th>S3 mAP@0.5 (End Point)</th>
<th>mAP &#x0394; (Robustness)</th>
<th>S3 F1-Score</th>
<th>S3 Recall (R)</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv12n</td>
<td>0.942</td>
<td>0.923</td>
<td>&#x2212;0.019</td>
<td>0.886</td>
<td>0.877</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>0.935</td>
<td>0.917</td>
<td>&#x2212;0.022</td>
<td>0.873</td>
<td>0.854</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>0.908</td>
<td>0.877</td>
<td>&#x2212;0.031</td>
<td>0.810</td>
<td>0.819</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>0.906</td>
<td>0.673</td>
<td>&#x2212;0.233</td>
<td>0.641</td>
<td>0.606</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>YOLOv10n suffered a catastrophic failure in this stage, with its mAP@0.5 plummeting by &#x2212;0.233. This strongly suggests a fundamental incompatibility between its architecture (possibly related to its NMS-free design) and the IFT strategy (especially when freezing the backbone).</p></list-item>
<list-item>
<p>YOLOv12n again demonstrated the best robustness (mAP dropped by only &#x2212;0.019) and ultimately completed the test with the highest mAP@0.5 (0.923) and F1-Score (0.886), solidifying its leadership across the entire IFT pipeline.</p></list-item>
</list></p>
<table-wrap id="table-16">
<label>Table A6</label>
<caption>
<title>S3 stage class performance breakdown (Macaque vs. laser spot)</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Class</th>
<th>S3 mAP@0.5</th>
<th>S3 Recall (R)</th>
<th>S3 F1-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">YOLOv12n</td>
<td>Macaque Monkey</td>
<td>0.910</td>
<td>0.838</td>
<td>0.861</td>
</tr>
<tr>
<td>Laser Spot</td>
<td>0.939</td>
<td>0.943</td>
<td>0.887</td>
</tr>
<tr>
<td rowspan="2">YOLOv11n</td>
<td>Macaque Monkey</td>
<td>0.917</td>
<td>0.835</td>
<td>0.861</td>
</tr>
<tr>
<td>Laser Spot</td>
<td>0.909</td>
<td>0.806</td>
<td>0.840</td>
</tr>
<tr>
<td rowspan="2">YOLOv10n</td>
<td>Macaque Monkey</td>
<td>0.856</td>
<td>0.721</td>
<td>0.825</td>
</tr>
<tr>
<td>Laser Spot</td>
<td>0.635</td>
<td>0.486</td>
<td>0.613</td>
</tr>
<tr>
<td rowspan="2">YOLOv8n</td>
<td>Macaque Monkey</td>
<td>0.903</td>
<td>0.880</td>
<td>0.862</td>
</tr>
<tr>
<td>Laser Spot</td>
<td>0.915</td>
<td>0.837</td>
<td>0.883</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>YOLOv12n and YOLOv11n both achieved an F1-Score of 0.861 for the old &#x201C;Macaque&#x201D; class, indicating the IFT strategy successfully preserved existing knowledge and avoided catastrophic forgetting.</p></list-item>
<list-item>
<p>YOLOv12n demonstrated a dominant advantage in learning the new &#x201C;Laser Spot&#x201D; class, with mAP@0.5 reaching 0.939 and an exceptional Recall of 0.943. This proves its feature extractor and A2C2f module are most effective for detecting instantaneous light sources and small targets.</p></list-item>
<list-item>
<p>YOLOv10n&#x2019;s catastrophic failure is primarily attributed to its extreme inability to handle the new class, with its Recall for &#x201C;Laser Spot&#x201D; being only 0.486.</p></list-item>
</list></p>

<table-wrap id="table-17">
<label>Table A7</label>
<caption>
<title>Comparison of comprehensive performance and edge deployment efficiency</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Final IFTmAP@0.5 (S3)</th>
<th>Final IFTF1-Score (S3)</th>
<th>Parameters (M)</th>
<th>Theoretical Complexity (GFLOPs)</th>
<th>FPS (Jetson AGX Orin, FP16)</th>
</tr>
</thead>
<tbody>
<tr>
<td>YOLOv12n</td>
<td>0.923</td>
<td>0.886</td>
<td>2.56M</td>
<td>6.3</td>
<td>480</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>0.917</td>
<td>0.873</td>
<td>2.58M</td>
<td>6.3</td>
<td>445</td>
</tr>
<tr>
<td>YOLOv8n</td>
<td>0.877</td>
<td>0.810</td>
<td>3.01M</td>
<td>8.1</td>
<td>383</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>0.673</td>
<td>0.641</td>
<td>2.27M</td>
<td>6.5</td>
<td>412</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Key Findings and Analysis:
<list list-type="bullet">
<list-item>
<p>YOLOv12n consistently demonstrated the best performance across both critical stages of the IFT process (S2 and S3), ultimately achieving leading mAP@0.5 (0.923) and F1-Score (0.886).</p></list-item>
<list-item>
<p>With its high signal-to-noise ratio pseudo-label retention rate of 47.1%, YOLOv12n was the only model to benefit (mAP gain of &#x002B;0.010) from semi-supervised learning.</p></list-item>
<list-item>
<p>YOLOv12n boasts the fewest parameters (2.56M) and a theoretical complexity of just 6.3 GFLOPs, representing a computational load reduction of approximately 22.22% compared to the SOTA baseline YOLOv8n (8.1 GFLOPs).</p></list-item>
<list-item>
<p>YOLOv12n achieved a high recall rate of 0.943 for the newly added &#x201C;laser spot&#x201D; class, proving its superior stability in reacting to high-speed flickering or weak light signals, thus making it highly suitable for the &#x201C;dynamic tracking&#x201D; task scenario of this research.</p></list-item>
</list></p>
<table-wrap id="table-18">
<label>Table A8</label>
<caption>
<title>Model weights and RMSE empirical values at different calibration point counts</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Statistical metric</th>
<th>N &#x003D; 50 Calibration points</th>
<th>N &#x003D; 100 Calibration points</th>
<th>N &#x003D; 181 Calibration points (Final Adopted)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Estimated weights (<inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>)</td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td><inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">0</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> (Intercept)</td>
<td>97.989</td>
<td>98.250</td>
<td>98.527</td>
</tr>
<tr>
<td><inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> (Linear Coeff.)</td>
<td>&#x2212;0.05156</td>
<td>&#x2212;0.05105</td>
<td>&#x2212;0.05074</td>
</tr>
<tr>
<td><inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> (Quadratic Coeff.)</td>
<td>0.000015</td>
<td>0.000013</td>
<td>0.000012</td>
</tr>
<tr>
<td colspan="4" align="center">Weight Estimation Standard Error (SE)</td>
</tr>
<tr>
<td><inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mrow><mml:mi mathvariant="bold-italic">S</mml:mi><mml:mi mathvariant="bold-italic">E</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">0</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td>0.52</td>
<td>0.35</td>
<td>0.28</td>
</tr>
<tr>
<td><inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mrow><mml:mi mathvariant="bold-italic">S</mml:mi><mml:mi mathvariant="bold-italic">E</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td>0.0008</td>
<td>0.0005</td>
<td>0.0004</td>
</tr>
<tr>
<td><inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mrow><mml:mi mathvariant="bold-italic">S</mml:mi><mml:mi mathvariant="bold-italic">E</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td>0.0000025</td>
<td>0.0000018</td>
<td>0.0000015</td>
</tr>
<tr>
<td>Final Measured RMSE (pixels)</td>
<td>1.8 pixels</td>
<td>1.4 pixels</td>
<td>1.2 pixels</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-19">
<label>Table A9</label>
<caption>
<title>Perturbation analysis empirical data</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Weight<break/>perturbation</th>
<th>Perturbation value (% of Original Coeff.)</th>
<th>Perturbed RMSE (pixels)</th>
<th>Observation and conclusion</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">0</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula></td>
<td>&#x00B1;1% (98.527 &#x00B1; 0.985)</td>
<td>1.5</td>
<td>RMSE increases, indicating the intercept is crucial for baseline offset calibration of the overall aiming position.</td>
</tr>
<tr>
<td></td>
<td>&#x00B1;5% (98.527 &#x00B1; 4.926)</td>
<td>2.5</td>
<td>A significant increase, indicating that even minor baseline offset errors can severely impact overall targeting accuracy.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">1</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula></td>
<td>&#x00B1;1% (&#x2212;0.05074 &#x00B1; 0.00051)</td>
<td>1.8</td>
<td>RMSE increases, confirming that the precision of <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mrow><mml:mover><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is decisive for the overall linear mapping accuracy of the calibration curve.</td>
</tr>
<tr>
<td></td>
<td>&#x00B1;5% (&#x2212;0.05074 &#x00B1; 0.00254)</td>
<td>3.8</td>
<td>RMSE increases rapidly, indicating that even a slight deviation in the laser angle-to-pixel scale can lead to significant aiming errors.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mrow><mml:mover><mml:msub><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula></td>
<td>&#x00B1;1% (0.000012 &#x00B1; 0.00000012)</td>
<td>1.6</td>
<td>RMSE increases, especially more pronounced at the image edges. This result demonstrates the critical role of the quadratic term in correcting non-linear distortions.</td>
</tr>
<tr>
<td></td>
<td>&#x00B1;5% (0.000012 &#x00B1; 0.0000006)</td>
<td>3.0</td>
<td>Even if the absolute value of the quadratic term coefficient is very small, its perturbation still leads to a significant increase in RMSE. This result clearly demonstrates the critical role of the quadratic term in correcting non-linear distortions, especially in areas far from the image center, proving its indispensability in the model.</td>
</tr>
</tbody>
</table>
</table-wrap>

<table-wrap id="table-20">
<label>Table A10</label>
<caption>
<title>Comparison of intra-class cosine distance distribution peak for ByteTrack/Bot-SORT</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Tracker</th>
<th>Intra-class cosine distance distribution peak (Average)</th>
<th>Identification capability conclusion</th>
</tr>
</thead>
<tbody>
<tr>
<td>ByteTrack</td>
<td>Highly concentrated near 0.1</td>
<td>Feature distinction capability is severely insufficient; appearance feature vectors of different monkeys are extremely similar.</td>
</tr>
<tr>
<td>Bot-SORT</td>
<td>Concentrated in the range of 0.4 to 0.5</td>
<td>Possesses excellent feature distinction capability, generating significantly different feature vectors for different individuals.</td>
</tr>
</tbody>
</table>
</table-wrap>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Na</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jung</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JY</given-names></string-name></person-group>. <article-title>Protection of orchard from wild animals and birds using USN facilities</article-title>. In: <conf-name>Proceedings of the 2nd International Conference on Computer and Automation Engineering (ICCAE); 2010 Feb 26&#x2013;28</conf-name>; <publisher-loc>Singapore</publisher-loc>. p. <fpage>307</fpage>&#x2013;<lpage>11</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hollinshead</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Briskie</surname> <given-names>JV</given-names></string-name>, <string-name><surname>Kross</surname> <given-names>SM</given-names></string-name></person-group>. <article-title>Orchard management factors affecting rates of bud damage to kiwifruit orchards</article-title>. <source>Crop Prot</source>. <year>2024</year>;<volume>184</volume>:<fpage>106792</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cropro.2024.106792</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koirala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Garber</surname> <given-names>PA</given-names></string-name>, <string-name><surname>Somasundaram</surname> <given-names>D</given-names></string-name>, <string-name><surname>Katuwal</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>B</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Factors affecting the crop raiding behavior of wild rhesus macaques in Nepal: implications for wildlife management</article-title>. <source>J Environ Manage</source>. <year>2021</year>;<volume>297</volume>:<fpage>113331</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jenvman.2021.113331</pub-id>; <pub-id pub-id-type="pmid">34298347</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Besala</surname> <given-names>FI</given-names></string-name>, <string-name><surname>Niimoto</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Okamoto</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Development of AI-based smart box trap system for capturing a harmful wild boar</article-title>. <source>ROBOMECH J</source>. <year>2025</year>;<volume>12</volume>(<issue>1</issue>):<fpage>4</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40648-025-00290-w</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Gross</surname> <given-names>EM</given-names></string-name>, <string-name><surname>Jayasinghe</surname> <given-names>N</given-names></string-name>, <string-name><surname>Brooks</surname> <given-names>A</given-names></string-name>, <string-name><surname>Polet</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wadhwa</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hilderink-Koopmans</surname> <given-names>F</given-names></string-name></person-group>. <source>A future for all: the need for human-wildlife coexistence</source>. <publisher-loc>Gland, Switzerland</publisher-loc>: <publisher-name>United Nations Environment Programme (UNEP) &#x0026; World Wide Fund for Nature (WWF)</publisher-name>; <year>2021</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>H</given-names></string-name></person-group>. <source>Human-macaque interactions and conflict: mitigating conflicting between humans and Taiwanese macaques (Macaca cyclopis) in Taiwan/Primates in Asian anthropogenic environments</source>. <publisher-loc>Tokyo, Japan</publisher-loc>: <publisher-name>Tokyo University of Foreign Studies</publisher-name>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ooko</surname> <given-names>SO</given-names></string-name>, <string-name><surname>Ndashimye</surname> <given-names>E</given-names></string-name>, <string-name><surname>Twahirwa</surname> <given-names>E</given-names></string-name>, <string-name><surname>Busogi</surname> <given-names>M</given-names></string-name></person-group>. <article-title>IoT and machine learning for smart bird monitoring and repellence: techniques, challenges, and opportunities</article-title>. <source>IoT</source>. <year>2025</year>;<volume>6</volume>(<issue>3</issue>):<fpage>46</fpage>. doi:<pub-id pub-id-type="doi">10.3390/iot6030046</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shim</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A deep learning-based intelligent wild boar repellent device</article-title>. <source>J Korean Inst Inf Technol</source>. <year>2021</year>;<volume>19</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.14801/jkiit.2021.19.5.1</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Afridi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Laporte-Devylder</surname> <given-names>L</given-names></string-name>, <string-name><surname>Maalouf</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kline</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Penny</surname> <given-names>SG</given-names></string-name>, <string-name><surname>Hlebowicz</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Impact of drone disturbances on wildlife: a review</article-title>. <source>Drones</source>. <year>2025</year>;<volume>9</volume>(<issue>4</issue>):<fpage>311</fpage>. doi:<pub-id pub-id-type="doi">10.3390/drones9040311</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Alaska Department of Transportation and Public Facilities</collab></person-group>. <article-title>Fairbanks international airport welcomes aurora to state service, Alaska&#x2019;s robotic solution to reducing airport wildlife conflicts [Internet]</article-title>. <year>2024 [cited 2025 Oct 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://dot.alaska.gov/faiiap/pdfs/PRs/072424-press-release.shtml">https://dot.alaska.gov/faiiap/pdfs/PRs/072424-press-release.shtml</ext-link>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mishra</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yadav</surname> <given-names>KK</given-names></string-name></person-group>. <article-title>Smart animal repelling device: utilizing IoT and AI for effective anti-adaptive harmful animal deterrence</article-title>. <source>BIO Web Conf</source>. <year>2024</year>;<volume>82</volume>(<issue>2</issue>):<fpage>05014</fpage>. doi:<pub-id pub-id-type="doi">10.1051/bioconf/20248205014</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A forest wildlife detection algorithm based on improved YOLOv5s</article-title>. <source>Animals</source>. <year>2023</year>;<volume>13</volume>(<issue>19</issue>):<fpage>3134</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ani13193134</pub-id>; <pub-id pub-id-type="pmid">37835740</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hui</surname> <given-names>X</given-names></string-name>, <string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Research on improved YOLOv5 for low-light environment object detection</article-title>. <source>Electronics</source>. <year>2023</year>;<volume>12</volume>(<issue>14</issue>):<fpage>3089</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics12143089</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>G</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Image-adaptive YOLO for object detection in adverse weather conditions</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2022</year>;<volume>36</volume>(<issue>2</issue>):<fpage>1792</fpage>&#x2013;<lpage>800</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v36i2.20072</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Video object tracking based on YOLOv7 and DeepSORT</article-title>. <comment>arXiv:2207.12202</comment>. <year>2022</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>E</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Detection of <italic>Camellia oleifera</italic> fruit in complex scenes by using YOLOv7 and data augmentation</article-title>. <source>Appl Sci</source>. <year>2022</year>;<volume>12</volume>(<issue>22</issue>):<fpage>11318</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app122211318</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chappidi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sundaram</surname> <given-names>DM</given-names></string-name></person-group>. <article-title>Novel animal detection system: cascaded YOLOv8 with adaptive preprocessing and feature extraction</article-title>. <source>IEEE Access</source>. <year>2024</year>;<volume>12</volume>:<fpage>110575</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2024.3439230</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Ultralytics</collab></person-group>. <article-title>UltralyticsYOLO12 (version 12.0.0) [Internet]</article-title>. <year>2025 [cited 2025 Oct 1]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/ultralytics/ultralytics">https://github.com/ultralytics/ultralytics</ext-link>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ogawa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yachida</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hosoi</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Multi object tracking based on uncertainty-aware RE-ID</article-title>. In: <conf-name>Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP); 2022 Oct 16&#x2013;19</conf-name>; <publisher-loc>Bordeaux, France</publisher-loc>. p. <fpage>346</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wojke</surname> <given-names>N</given-names></string-name>, <string-name><surname>Bewley</surname> <given-names>A</given-names></string-name>, <string-name><surname>Paulus</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Simple online and realtime tracking with a deep association metric</article-title>. In: <conf-name>Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP); 2017 Sep 17&#x2013;20</conf-name>; <publisher-loc>Beijing, China</publisher-loc>. p. <fpage>3645</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>P</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Bytetrack: multi-object tracking by associating every detection box</article-title>. In: <conf-name>Proceedings of the Computer Vision&#x2014;ECCV 2022; 2022 Oct 23&#x2013;27</conf-name>; <publisher-loc>Tel Aviv, Israel</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Observation-centric SORT: rethinking SORT for robust multi-object tracking</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 18&#x2013;22</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>9686</fpage>&#x2013;<lpage>96</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Aharon</surname> <given-names>N</given-names></string-name>, <string-name><surname>Orfaig</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bobrovsky</surname> <given-names>BZ</given-names></string-name></person-group>. <article-title>BoT-SORT: robust associations multi-pedestrian tracking</article-title>. <comment>arXiv:2206.14651</comment>. <year>2022</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ryan</surname> <given-names>SZ</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>R</given-names></string-name>, <string-name><surname>Joppa</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Reinforcement learning for green security games with real-time information</article-title>. <source>Proc 38th AAAI Conf Artif Intell</source>. <year>2019</year>;<volume>33</volume>(<issue>1</issue>):<fpage>1401</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v33i01.33011401</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luongo</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>CLA</given-names></string-name>, <string-name><surname>Hesse</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Wekselblatt</surname> <given-names>JB</given-names></string-name>, <string-name><surname>Lanfranchi</surname> <given-names>FF</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mice and Primates use distinct strategies for visual segmentation</article-title>. <source>eLife</source>. <year>2023</year>;<volume>12</volume>:<fpage>e74394</fpage>. doi:<pub-id pub-id-type="doi">10.7554/eLife.74394</pub-id>; <pub-id pub-id-type="pmid">36790170</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duporge</surname> <given-names>I</given-names></string-name>, <string-name><surname>Kholiavchenko</surname> <given-names>M</given-names></string-name>, <string-name><surname>Harel</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wolf</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rubenstein</surname> <given-names>DI</given-names></string-name>, <string-name><surname>Crofoot</surname> <given-names>MC</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>BaboonLand dataset: tracking primates in the wild and automating behavior recognition from drone videos</article-title>. <source>Int J Comput Vis</source>. <year>2025</year>;<volume>133</volume>:<fpage>6578</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-025-02532-1</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Research on cattle behavior recognition and multi-object tracking</article-title>. <source>Animals</source>. <year>2024</year>;<volume>14</volume>(<issue>2993</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.3390/ani14202993</pub-id>; <pub-id pub-id-type="pmid">39457923</pub-id></mixed-citation></ref>
</ref-list>
</back></article>