<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">83923</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.083923</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>BMGKD: A High Precision Object Detection Knowledge Distillation Method for Bridging Multi-Dimensional Gaps</article-title>
<alt-title alt-title-type="left-running-head">BMGKD: A High Precision Object Detection Knowledge Distillation Method for Bridging Multi-Dimensional Gaps</alt-title>
<alt-title alt-title-type="right-running-head">BMGKD: A High Precision Object Detection Knowledge Distillation Method for Bridging Multi-Dimensional Gaps</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Tianqi</given-names></name></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Yang</given-names></name></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Pan</surname><given-names>Zhisong</given-names></name><email>panzhisong@aeu.edu.cn</email></contrib>
<aff id="aff-1">
<institution>College of Command and Control Engineering, Army Engineering University of PLA</institution>, <addr-line>Nanjing</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Zhisong Pan. Email: <email>panzhisong@aeu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>07</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>3</issue>
<elocation-id>81</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>04</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>06</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_83923.pdf"></self-uri>
<abstract>
<p>Existing knowledge distillation methods for object detection struggle to bridge the teacher-student capacity gap and overlook the inherent differences between classification and regression subtasks. To address these issues, we propose a Bridging Multi-dimensional Gaps Knowledge Distillation (BMGKD) method, which comprises two core modules: a feature difference distillation module and a response difference distillation module. The feature difference distillation module achieves global feature structural alignment via improved centered kernel alignment and performs local key feature alignment using joint spatial and channel-wise cosine similarity masks. The response difference distillation module constructs a dynamic classification mask and a high-quality prediction box selection mechanism, along with a classification and regression co-optimization loss function. On the MS COCO dataset, BMGKD achieves the highest mAP across all three teacher-student configurations on two-stage Faster R-CNN, single-stage anchor-free GFL, and single-stage anchor-based RetinaNet detectors. In the ResNet101<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>ResNet50 setting, it surpasses baselines by 2.7, 4.2, and 2.8 percentage points, respectively, and maintains consistent optimality under varying capacity gaps and cross-architecture scenarios. On Pascal VOC 2007, BMGKD attains mAP of 56.2<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, 58.3<inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and 58.0<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on the same three detectors, outperforming all compared methods. Ablation studies confirm each component&#x2019;s contribution. These results demonstrate that BMGKD effectively resolves the distillation accuracy degradation caused by capacity gaps and task discrepancies between teacher and student models. It provides an effective solution for knowledge distillation across diverse detection architectures.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Object detection</kwd>
<kwd>knowledge distillation</kwd>
<kwd>model compression</kwd>
<kwd>feature fusion</kwd>
</kwd-group><funding-group>
<award-group id="awg1">
<funding-source>Army Engineering University of PLA</funding-source>
</award-group>
</funding-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, deep learning has achieved significant progress and remarkable results in various applications, particularly in visual tasks such as image classification [<xref ref-type="bibr" rid="ref-1">1</xref>] and object detection [<xref ref-type="bibr" rid="ref-2">2</xref>]. Benefiting from the superior performance of deep learning models, existing relevant methods have been widely integrated and deployed in practical scenarios, including autonomous driving, smart retail, and drone operations. Although deep learning models have demonstrated excellent performance, their deployment on mobile and edge devices remains a major challenge, primarily due to the severe constraints on computing power and storage resources of such devices. To address this issue, knowledge distillation (KD) [<xref ref-type="bibr" rid="ref-3">3</xref>] has been proposed as a solution. It effectively reduces the parameter count and computational overhead of deep convolutional neural networks (CNNs) while maximizing the preservation of model performance.</p>
<p>Knowledge distillation was initially applied to the lightweight compression of classification models, and numerous researchers have since extended it to the field of object detection. Li et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] directly selected features from the teacher model as distillation supervision targets. Building upon the soft labels introduced in [<xref ref-type="bibr" rid="ref-3">3</xref>], Chen et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] directly applied them to the classification branch of object detection models. Wu et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] designed refined feature masks to enable feature imitation learning between teacher and student models. Lan and Tian [<xref ref-type="bibr" rid="ref-7">7</xref>] utilized gradient information to filter and distill high-contribution features. Despite the significant progress made by the aforementioned works in object detection knowledge distillation, none of these studies have thoroughly explored the model capacity differences between teacher and student networks, nor have they conducted targeted research on the distinctions between the classification and regression tasks in object detection.</p>
<p>Regarding the intermediate feature layers of object detection models, existing feature mask designs rely excessively on the teacher model. These masks are typically generated by extracting feature maps from the intermediate layers of the teacher network and performing simple attention calculations [<xref ref-type="bibr" rid="ref-8">8</xref>]. This approach has two major limitations. First, mask generation relies solely on the teacher model&#x2019;s perspective, completely ignoring the student model&#x2019;s learning state. Second, attention is extracted only based on the basic information of feature maps, failing to uncover their deep-level correlations. Consequently, during training, the student model cannot fully perceive and understand the semantic meaning of key features filtered by the teacher model through masking. Furthermore, regarding the final output layer of object detection models, traditional knowledge distillation algorithms [<xref ref-type="bibr" rid="ref-9">9</xref>] use kullback-leibler (KL) divergence to uniformly align the outputs of the classification and regression heads, failing to account for the distinct characteristics of these two subtasks. This leads to the underutilization of a substantial amount of valuable output information, ultimately resulting in limited distillation performance.</p>
<p>To address these issues, we propose a knowledge distillation method for object detection that bridges multidimensional gaps between teacher and student models. BMGKD constructs a feature difference module that synchronizes the correlation between the global and local feature maps of the teacher and student models in key regions. This enhances the student model&#x2019;s ability to perceive and understand features, thereby bridging the performance gap between teacher and student networks. Additionally, we design a response difference module to achieve the joint optimization of classification and regression in object detection tasks. Experimental results on the MS COCO dataset demonstrate that BMGKD achieves a detection accuracy of 44.4<inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on a single-stage anchor-free detector, representing a 4.2<inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> improvement over the baseline model.</p>
<p>Our main contributions are summarized as follows:</p>
<p>(1) We propose the BMGKD framework, which bridges feature discrepancies through global alignment via the Gram matrix and local alignment via space-channel joint masking, and bridges response discrepancies through dynamic classification masking and region selection.</p>
<p>(2) Within the response difference distillation module, we establish a collaborative optimization mechanism for classification and regression, implementing joint teacher-student decision-making and a task-consistency loss design.</p>
<p>(3) Across three detectors, multiple capacity gaps, and two datasets, BMGKD outperforms over a dozen mainstream distillation methods from recent years, with ablation studies and visualizations confirming the effectiveness of each component.</p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> provides an overview of related work. The details of BMGKD are described in <xref ref-type="sec" rid="s3">Section 3</xref>. In <xref ref-type="sec" rid="s4">Section 4</xref>, two public datasets are used to evaluate the performance of BMGKD. <xref ref-type="sec" rid="s5">Section 5</xref> provides a summary of the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Object Detection</title>
<p>Object detection is one of the most fundamental and actively researched topics in the field of computer vision, with significant practical value in industrial applications. Based on algorithmic architecture and detection processes, existing object detectors can generally be categorized into three types: two-stage detectors, single-stage anchor-based detectors, and single-stage anchor-free detectors. Two-stage detectors use neural networks to extract candidate regions and then perform precise classification and bounding box regression on each candidate region. Representative algorithms of this class of detectors include R-CNN and its improved version, Fast R-CNN [<xref ref-type="bibr" rid="ref-10">10</xref>]. Among these, Faster R-CNN achieved end-to-end training by introducing a region proposal network(RPN). Building on this foundation, various RPN-based improved algorithms have emerged, including Cascaded R-CNN and Mask R-CNN [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
<p>Two-stage detectors achieve higher detection accuracy but have slower inference speeds. Compared with two-stage detectors, single-stage detectors can simultaneously perform object classification and regression tasks, significantly improving inference speed but at the cost of reduced detection accuracy. Single-stage anchor-based detectors, such as RetinaNet [<xref ref-type="bibr" rid="ref-12">12</xref>] and YOLO [<xref ref-type="bibr" rid="ref-13">13</xref>] complete detection tasks using predefined anchor boxes, but this introduces additional hyperparameters and requires redundant anchor box pruning mechanisms during the processing. In contrast, single-stage anchor-free detectors, represented by FCOS [<xref ref-type="bibr" rid="ref-14">14</xref>] and GFL [<xref ref-type="bibr" rid="ref-15">15</xref>], abandon the anchor box design and use keypoint regression to directly predict bounding boxes.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Knowledge Distillation</title>
<p>Knowledge distillation is an efficient model compression technique that extracts and transfers implicit supervisory knowledge from a high performance teacher model to train a lightweight student model. It enables the student model to achieve performance close to that of the teacher model while maintaining its lightweight nature. Existing methods are primarily divided into two categories: feature-based distillation and response-based distillation.</p>
<p>Feature-based distillation focuses on the multi-scale features output by feature pyramid networks, extracting the implicit knowledge embedded in these intermediate-layer features. As an early knowledge distillation model targeting intermediate features, FitNet [<xref ref-type="bibr" rid="ref-16">16</xref>] demonstrated that the intermediate-layer features of the teacher network contain rich local structural and pattern information. This information can serve as an effective supervisory signal to guide the training of the student model. Lv et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] introduced feature map transformation and alignment mechanisms to optimize the distribution differences between teacher and student features. Yang et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] proposed a feature masking knowledge distillation method that can adaptively select key regions from the teacher&#x2019;s feature maps. Li et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] proposed CSAKD, which utilized the teacher&#x2019;s attention information to generate spatial masks and enhanced student features by adding teacher global context. However, CSAKD relied on the teacher&#x2019;s unilateral attention to determine the masking regions, ignoring the real time learning state and capability gap of the student model.</p>
<p>Response-based distillation focuses on the final network outputs generated by classification and regression tasks. The LD method [<xref ref-type="bibr" rid="ref-19">19</xref>] adopted the bounding box probability distribution proposed by GFL, used the KL divergence to align the outputs of the teacher and student models. It also designed a region-weighting strategy based on the intersection over union (IoU) ratio to optimize knowledge distillation performance. BCKD [<xref ref-type="bibr" rid="ref-20">20</xref>] reconstructed the classification branch output logic into multiple sets of binary classification results, effectively alleviating the mismatch between detection tasks and classification loss functions. Zhang et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed SAMKD, which introduced a spatial pyramid mechanism and used the difference between teacher and student features to weight logit distillation. Although SAMKD achieved hierarchical logit alignment at multiple scales, it uniformly applied KL divergence for logit distillation across all levels, thereby overlooking the fundamental task disparity between classification and regression heads in object detection.</p>
<p>However, feature-based distillation methods fail to fully exploit the capacity differences between teacher and student models, resulting in insufficient alignment between feature supervision and the student model. Response-based distillation methods overlook the task differences between the classification head and the regression head in detection tasks. In contrast, the method proposed in this paper further refines the features of the neck network and the final response outputs of the object detector by introducing feature difference and response difference modules.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<sec id="s3_1">
<label>3.1</label>
<title>Overall Structure</title>
<p>The proposed BMGKD framework consists of a feature difference distillation module and a response difference distillation module. During the model training phase, both the teacher and student networks perform forward passes simultaneously. The parameters of the teacher network are fully frozen, while the student network learns by transferring knowledge from the teacher network and utilizing supervision from the detection task, iteratively updating its own parameters through backpropagation. Subsequently, feature maps are extracted from each layer of the teacher-student feature pyramid network and processed through the feature difference module to achieve dual alignment of global and local features. Meanwhile, the response difference module selectively distills response values from the teacher and student network branches, focusing on optimizing regions where the outputs of the teacher and student networks differ significantly in classification and regression. Additionally, we designed a dedicated loss function to bridge the task gap between the classification head and the regression head. The overall network architecture of BMGKD is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overall network architecture of BMGKD.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83923-fig-1.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Feature Difference Distillation Module</title>
<p>To ensure that the student model can fully learn from the teacher model, we propose a feature difference distillation module. This module uses a Gram matrix-based CKA loss function to distill the structural relationships between features rather than their absolute values. Furthermore, by combining spatial and channel-wise cosine similarity masks, it dynamically assigns weights to each distillation region based on the differences between the teacher and student models. This achieves comprehensive alignment between the global and local features of the teacher and student models, thereby improving the student model&#x2019;s learning efficiency and its ability to understand features.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Global Feature Difference Alignment</title>
<p>In the global feature processing stage, the feature difference distillation module does not perform direct numerical alignment on the global feature map. Instead, it introduces Gram matrices to model the relationships between features.</p>
<p>Let the teacher model be denoted by <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>t</mml:mi></mml:math></inline-formula> and the lightweight student model by <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>s</mml:mi></mml:math></inline-formula>. The feature map output by the <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>i</mml:mi></mml:math></inline-formula>-th layer of the teacher model&#x2019;s neck network is denoted by <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>h</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>w</mml:mi></mml:math></inline-formula> represent the number of channels, height, and width of the teacher model&#x2019;s feature map, respectively. The feature map output by the <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>i</mml:mi></mml:math></inline-formula>-th layer of the student model&#x2019;s neck network is denoted by <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>h</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>w</mml:mi></mml:math></inline-formula> represent the number of channels, height, and width of the student model&#x2019;s feature map, respectively. <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>b</mml:mi></mml:math></inline-formula> is the training batch size. Use <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to transform <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> into <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>T</mml:mi></mml:math></inline-formula> is transformation function used to unify the dimensions of the teacher and student feature maps. The CKA method is effective for knowledge distillation tasks [<xref ref-type="bibr" rid="ref-22">22</xref>], but the computational complexity of the traditional CKA is high. Therefore, we have improved and optimized it, as shown in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>K</mml:mi><mml:mi>A</mml:mi></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:msup><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>X</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>X</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>Y</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mi>Y</mml:mi><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> are Gram-Matrix. <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is a vectorization operation. For the feature matrix <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, the similarity between the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>i</mml:mi></mml:math></inline-formula>-th and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>j</mml:mi></mml:math></inline-formula>-th samples in the feature space of its Gram matrix <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>G</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is denoted by <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>. As can be seen from the formula, the improved CKA metric is equivalent to the cosine similarity of teacher-student features after vectorizing the Gram matrix.</p>
<p>The Gram matrix elevates traditional eigenvalue matching to the level of feature relationship structure matching, enabling the student network to simultaneously learn both the teacher&#x2019;s feature representations and the logic of feature organization. The distillation loss function used for global feature alignment is shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>Global</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:msup><mml:mi>X</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:msup><mml:mi>Y</mml:mi><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Local Feature Difference Alignment</title>
<p>In the phase of aligning local feature differences, to ensure that the student model focuses on the most discriminative feature regions in the teacher model, we use the extracted teacher and student feature maps to calculate spatial cosine similarity and channel cosine similarity, respectively. We then construct a local key information mask based on these two metrics.</p>
<p>First, the feature mapping layer <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is used to transform the dimensions of the student network&#x2019;s feature map so that they match the number of channels in the teacher network. Subsequently, the spatial cosine similarity <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msup><mml:mi>C</mml:mi><mml:mi>S</mml:mi></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and channel cosine similarity <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msup><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are calculated for the teacher feature maps and student feature maps at the corresponding levels, respectively. For each spatial location <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> on the feature map, the teacher and student channel vectors at that location are defined as follows:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The spatial cosine similarity <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup></mml:math></inline-formula> is the cosine similarity between the vectors of these two channels:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:msqrt><mml:mo>&#x22C5;</mml:mo><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>C</mml:mi></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the value of the teacher feature map at channel <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>k</mml:mi></mml:math></inline-formula>-th and position <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the value of the student feature map at the corresponding position. For each channel <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>k</mml:mi></mml:math></inline-formula>, flatten the spatial feature map on that channel into an <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:math></inline-formula> dimensional vector:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The channel cosine similarity <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msubsup><mml:mi>C</mml:mi><mml:mi>k</mml:mi><mml:mi>C</mml:mi></mml:msubsup></mml:math></inline-formula> is the cosine similarity between the vectors of these two spatial:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mi>C</mml:mi><mml:mi>k</mml:mi><mml:mi>C</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:msubsup></mml:mrow><mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">u</mml:mi></mml:mrow><mml:mi>k</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>h</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>w</mml:mi></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>h</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>w</mml:mi></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:msqrt><mml:mo>&#x22C5;</mml:mo><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>h</mml:mi></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>w</mml:mi></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>By combining the spatial cosine similarity <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> with the channel cosine similarity <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref> to construct a joint feature mask, we achieve precise screening of key regions in the teacher-student feature maps. The calculation formula is shown in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msubsup><mml:mi>M</mml:mi><mml:mi>J</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mi>S</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>k</mml:mi><mml:mi>C</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The joint mask, constructed based on cosine similarity, integrates prior knowledge of teacher features with the real-time learning state of the student model. It addresses the shortcoming of traditional methods, where region selection is solely teacher-driven, while enhancing the student model&#x2019;s understanding of the relationships among teacher features, thereby minimizing the differences between teacher and student feature maps. The final loss function for local feature differences is shown in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:msubsup><mml:mi>M</mml:mi><mml:mi>J</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Response Difference Distillation Module</title>
<p>Traditional knowledge distillation methods for object detection optimize the response outputs of the classification and regression heads independently. This results in a significant amount of useful information being overlooked in both branches, further degrading the distillation performance. To overcome this limitation, the response difference distillation module enables collaborative optimization and joint distillation of the classification and regression tasks. For the predicted scores output by the classification head, we define the base classification mask <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msup><mml:mi>M</mml:mi><mml:mi>t</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:munder><mml:mo form="prefix">max</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msubsup><mml:mi>p</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msubsup></mml:math></inline-formula> is the prediction score of the teacher model for the <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>c</mml:mi></mml:math></inline-formula>-th target class. This mask is a static mask and cannot be adapted during the model training process. Therefore, based on the design of the feature-differential distillation module, we construct a dynamic classification mask <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msup><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munder><mml:mo form="prefix">max</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mo>|</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>|</mml:mo></mml:mrow></mml:math></inline-formula> that employs joint teacher-student decision-making.</p>
<p>Based on the coordinate values output by the regression head, we use a region selection function to identify high quality prediction boxes and filter out low quality ones during the regression output refinement stage. <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msup><mml:mi>B</mml:mi><mml:mi>s</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is the set of prediction boxes for a single image in the student model, containing a total of <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>R</mml:mi></mml:math></inline-formula> prediction boxes. <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msup><mml:mi>B</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msubsup></mml:math></inline-formula> is the set of ground-truth bounding boxes corresponding to the image, containing a total of <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> ground-truth boxes. The region selection function for the <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>r</mml:mi></mml:math></inline-formula>-th prediction box <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup></mml:math></inline-formula> generated by the student model is defined as:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:munder><mml:mo form="prefix">max</mml:mo><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:mrow><mml:mtext>IoU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003E;</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mtext>IoU</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mrow><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the formula for calculating the standard merger ratio. <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>I</mml:mi></mml:math></inline-formula> is the indicator function. When the maximum intersection ratio exceeds the set threshold <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>t</mml:mi></mml:math></inline-formula>, the prediction box is considered a positive sample. Otherwise, it is considered a negative sample.</p>
<p>To ensure consistency between distillation and classification tasks, binary cross-entropy loss is used as the basis for designing the classification head distillation loss function. This minimizes potential inconsistencies arising from differences in loss functions, thereby achieving more efficient knowledge distillation. A joint masked distillation operation is performed on the classification head distillation loss function to resolve these discrepancies, as shown in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>BCE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> represents the classification distillation loss, and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>BCE</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> represents the binary cross-entropy loss. <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>N</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> is the normalization factor. <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msubsup></mml:math></inline-formula> represent the prediction scores for class <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>k</mml:mi></mml:math></inline-formula> at position <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>i</mml:mi></mml:math></inline-formula> in the teacher detector and student detector, respectively.</p>
<p>In addition, for the regression loss of the object detection head, we adopt the <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>G</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:math></inline-formula> loss function, which takes into account not only the position and size of the bounding box but also its shape. The formula for the <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>G</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:math></inline-formula> loss is shown in <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup></mml:math></inline-formula> represent the <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>r</mml:mi></mml:math></inline-formula>-th bounding box from the teacher detector and the student detector, respectively. <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>I</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the intersection area and union area of the two bounding boxes, respectively. <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>A</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the area of the smallest bounding box that contains the two rectangles. The regression head distillation loss function is given by <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>N</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mi>M</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>GIoU</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>b</mml:mi><mml:mi>r</mml:mi><mml:mi>s</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Overall Loss Function</title>
<p>The overall loss function of the method proposed in this paper consists of five components: <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the original detection task loss of the student model. <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the global feature alignment loss of the feature difference distillation module. <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the local feature alignment loss of the feature difference distillation module. <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the classification head distillation loss of the response difference distillation module. <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the regression head distillation loss of the response difference distillation module. The overall loss function is defined as shown in <xref ref-type="disp-formula" rid="eqn-15">Eq.(15)</xref>:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>loss</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>det</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>global</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>local</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> are hyperparameters used to balance the individual loss components. The overall distillation loss is calculated solely based on the feature maps and response outputs, without involving the internal structure of the detector. Consequently, the proposed method exhibits excellent adaptability and can be applied to various mainstream object detection architectures.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset and Experimental Environment Configuration</title>
<p>To verify the effectiveness and efficiency of the proposed method, We conducted experimental validation using the MS COCO object detection dataset [<xref ref-type="bibr" rid="ref-23">23</xref>] and Pascal VOC 2007 dataset [<xref ref-type="bibr" rid="ref-24">24</xref>]. The training phase utilized the trainval135k subset, comprising 115,000 labeled images. The validation phase employed the minival validation subset, containing 5000 images, to assess model performance during training. Finally, the MS COCO test-dev 2019 dataset, which contains 20,000 images, was used for evaluation. Pascal VOC 2007 dataset contains 4952 images covering 20 object categories, including people, animals, vehicles, and indoor objects, among others.</p>
<p>The model was trained using the stochastic gradient descent (SGD) optimizer throughout the entire process, with a momentum coefficient of 0.9 and a weight decay coefficient of 0.0001. BMGKD is compatible with various mainstream object detection architectures, including the two-stage detector Faster-RCNN, the anchor-based single-stage detectors RetinaNet and GFL, and the anchor-free single-stage detector FCOS. All methods were trained on four NVIDIA A100 GPUs for 12 training epochs. The evaluation adopted the standard COCO metric of average precision (AP), and reporting the mean average precision (mAP) values at intersection-over-union (IoU) thresholds of 0.5 and 0.75, as well as the AP values for small, medium, and large-scale targets. All results are based on three independent replicates and report the average performance. The experiments demonstrate that the improvements made to BMGKD are stable and reliable, with a standard deviation of mAP between runs of less than 0.2<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Comparison Experiments</title>
<p>As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, we conducted a comprehensive performance comparison between our proposed method and other state-of-the-art methods (TBD [<xref ref-type="bibr" rid="ref-25">25</xref>], SamKD [<xref ref-type="bibr" rid="ref-21">21</xref>], CSAKD [<xref ref-type="bibr" rid="ref-18">18</xref>], FGD [<xref ref-type="bibr" rid="ref-8">8</xref>], BCKD [<xref ref-type="bibr" rid="ref-20">20</xref>], DrKD [<xref ref-type="bibr" rid="ref-26">26</xref>], CrossKD [<xref ref-type="bibr" rid="ref-27">27</xref>], HMKD [<xref ref-type="bibr" rid="ref-28">28</xref>], MLD [<xref ref-type="bibr" rid="ref-29">29</xref>], and PCD [<xref ref-type="bibr" rid="ref-30">30</xref>]) on the COCO dataset. On the two-stage Faster R-CNN, single-stage anchor-free GFL, and single-stage anchor-based RetinaNet detector architectures, we evaluate the performance of BMGKD under three teacher-student configurations: the standard capacity gap ResNet101 to ResNet50, the larger capacity gap ResNet101 to ResNet18, and the cross-architecture ResNet50 to MobileNetV3. In all experiments, BMGKD achieved the highest mAP and outperformed other models across all evaluation metrics, including <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Specifically, in the ResNet101<inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>ResNet50 configuration, BMGKD achieved mAPs of 41.1<inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, 44.4<inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and 40.4<inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on the three detectors, respectively, representing improvements of 2.7, 4.2, and 2.8 percentage points over the baseline, significantly outperforming existing methods. When the capacity gap between teachers and students widened further from ResNet101 to ResNet18, BMGKD still achieved mAPs of 37.4<inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, 40.0<inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and 34.6<inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, respectively, maintaining a clear lead. In the more challenging cross-architecture scenario from ResNet50 to MobileNetV3, BMGKD ranked first with mAP scores of 36.9<inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, 39.4<inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and 34.8<inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. These results fully demonstrate that BMGKD can consistently handle different detection architectures, varying degrees of capacity differences, and cross-backbone family transfer requirements, showcasing exceptional consistency and generalization capabilities.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of distillation results across different detectors and teacher-student configurations.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Detector</th>
<th>Teacher</th>
<th>Student</th>
<th>Method</th>
<th>mAP</th>
<th><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>AP</mml:mi><mml:mi>S</mml:mi></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>AP</mml:mi><mml:mi>M</mml:mi></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>AP</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="15">Faster R-CNN</td>
<td rowspan="5">ResNet101</td>
<td rowspan="5">ResNet50</td>
<td>Baseline</td>
<td>38.4</td>
<td>59.0</td>
<td>42.0</td>
<td>21.5</td>
<td>42.1</td>
<td>50.3</td>
</tr>
<tr>



<td>TBD (2023)</td>
<td>39.8 (&#x002B;1.4)</td>
<td>57.2</td>
<td>40.1</td>
<td>21.9</td>
<td>42.3</td>
<td>50.6</td>
</tr>
<tr>



<td>SamKD (2025)</td>
<td>40.6 (&#x002B;2.2)</td>
<td>61.0</td>
<td>44.3</td>
<td>23.5</td>
<td>44.7</td>
<td>51.0</td>
</tr>
<tr>



<td>CSAKD (2025)</td>
<td>40.7 (&#x002B;2.3)</td>
<td>61.1</td>
<td>44.4</td>
<td>23.6</td>
<td>44.7</td>
<td>51.4</td>
</tr>
<tr>



<td>BMGKD</td>
<td>41.1 (&#x002B;2.7)</td>
<td>61.5</td>
<td>44.7</td>
<td>23.8</td>
<td>45.0</td>
<td>52.5</td>
</tr>
<tr>

<td rowspan="5">ResNet101</td>
<td rowspan="5">ResNet18</td>
<td>Baseline</td>
<td>34.5</td>
<td>54.6</td>
<td>37.2</td>
<td>19.2</td>
<td>36.8</td>
<td>45.2</td>
</tr>
<tr>



<td>TBD (2023)</td>
<td>35.4 (&#x002B;0.9)</td>
<td>55.7</td>
<td>38.1</td>
<td>20.4</td>
<td>37.9</td>
<td>46.7</td>
</tr>
<tr>



<td>SamKD (2025)</td>
<td>36.0 (&#x002B;1.5)</td>
<td>56.2</td>
<td>38.6</td>
<td>20.7</td>
<td>38.3</td>
<td>46.8</td>
</tr>
<tr>



<td>CSAKD (2025)</td>
<td>36.6 (&#x002B;2.1)</td>
<td>56.7</td>
<td>39.5</td>
<td>21.6</td>
<td>39.4</td>
<td>47.4</td>
</tr>
<tr>



<td>BMGKD</td>
<td>37.4 (&#x002B;2.9)</td>
<td>57.4</td>
<td>40.0</td>
<td>21.9</td>
<td>39.8</td>
<td>48.1</td>
</tr>
<tr>

<td rowspan="5">ResNet50</td>
<td rowspan="5">MobileNetV3</td>
<td>Baseline</td>
<td>34.3</td>
<td>54.3</td>
<td>37.0</td>
<td>18.9</td>
<td>36.5</td>
<td>45.0</td>
</tr>
<tr>



<td>FGD (2022)</td>
<td>35.4 (&#x002B;1.1)</td>
<td>54.8</td>
<td>37.4</td>
<td>19.4</td>
<td>38.6</td>
<td>46.3</td>
</tr>
<tr>



<td>BCKD (2023)</td>
<td>36.2 (&#x002B;1.9)</td>
<td>56.4</td>
<td>39.1</td>
<td>20.9</td>
<td>39.0</td>
<td>47.0</td>
</tr>
<tr>



<td>DrKD (2025)</td>
<td>36.6 (&#x002B;2.3)</td>
<td>56.5</td>
<td>39.4</td>
<td>21.1</td>
<td>39.3</td>
<td>47.4</td>
</tr>
<tr>



<td>BMGKD</td>
<td>36.9 (&#x002B;2.6)</td>
<td>57.0</td>
<td>39.9</td>
<td>21.5</td>
<td>39.6</td>
<td>47.6</td>
</tr>
<tr>
<td rowspan="17">GFL</td>
<td rowspan="6">ResNet101</td>
<td rowspan="6">ResNet50</td>
<td>Baseline</td>
<td>40.2</td>
<td>58.4</td>
<td>43.3</td>
<td>23.3</td>
<td>44.0</td>
<td>52.2</td>
</tr>
<tr>



<td>BCKD (2023)</td>
<td>42.1 (&#x002B;1.9)</td>
<td>60.2</td>
<td>45.5</td>
<td>25.3</td>
<td>46.5</td>
<td>55.6</td>
</tr>
<tr>



<td>CrossKD (2024)</td>
<td>42.4 (&#x002B;2.2)</td>
<td>59.4</td>
<td>45.0</td>
<td>24.3</td>
<td>45.6</td>
<td>53.4</td>
</tr>
<tr>



<td>HMKD (2026)</td>
<td>42.9 (&#x002B;2.7)</td>
<td>61.6</td>
<td>46.9</td>
<td>25.7</td>
<td>47.3</td>
<td>55.9</td>
</tr>
<tr>



<td>DrKD (2025)</td>
<td>43.2 (&#x002B;3.0)</td>
<td>61.3</td>
<td>46.8</td>
<td>26.3</td>
<td>47.8</td>
<td>56.3</td>
</tr>
<tr>



<td>BMGKD</td>
<td>44.4 (&#x002B;4.2)</td>
<td>62.7</td>
<td>48.3</td>
<td>27.4</td>
<td>49.0</td>
<td>56.8</td>
</tr>
<tr>

<td rowspan="6">ResNet101</td>
<td rowspan="6">ResNet18</td>
<td>Baseline</td>
<td>35.7</td>
<td>53.5</td>
<td>38.0</td>
<td>19.7</td>
<td>38.7</td>
<td>46.6</td>
</tr>
<tr>



<td>BCKD (2023)</td>
<td>37.8 (&#x002B;2.1)</td>
<td>55.7</td>
<td>40.3</td>
<td>21.8</td>
<td>40.9</td>
<td>48.9</td>
</tr>
<tr>



<td>CrossKD (2024)</td>
<td>38.4 (&#x002B;2.4)</td>
<td>56.1</td>
<td>40.7</td>
<td>22.0</td>
<td>41.2</td>
<td>49.4</td>
</tr>
<tr>



<td>HMKD (2026)</td>
<td>38.7 (&#x002B;3.0)</td>
<td>56.8</td>
<td>41.2</td>
<td>22.4</td>
<td>41.9</td>
<td>50.0</td>
</tr>
<tr>



<td>DrKD (2025)</td>
<td>38.6 (&#x002B;2.9)</td>
<td>56.8</td>
<td>41.3</td>
<td>22.4</td>
<td>42.0</td>
<td>49.8</td>
</tr>
<tr>



<td>BMGKD</td>
<td>40.0 (&#x002B;4.3)</td>
<td>57.9</td>
<td>42.3</td>
<td>23.5</td>
<td>43.2</td>
<td>51.4</td>
</tr>
<tr>

<td rowspan="5">ResNet50</td>
<td rowspan="5">MobileNetV3</td>
<td>Baseline</td>
<td>35.4</td>
<td>53.3</td>
<td>37.7</td>
<td>19.4</td>
<td>38.5</td>
<td>46.4</td>
</tr>
<tr>



<td>SamKD (2025)</td>
<td>38.0 (&#x002B;2.6)</td>
<td>56.1</td>
<td>40.4</td>
<td>22.1</td>
<td>41.0</td>
<td>49.2</td>
</tr>
<tr>



<td>DrKD (2025)</td>
<td>38.0 (&#x002B;2.6)</td>
<td>56.2</td>
<td>40.3</td>
<td>22.0</td>
<td>41.2</td>
<td>49.4</td>
</tr>
<tr>



<td>HMKD (2026)</td>
<td>38.7 (&#x002B;3.3)</td>
<td>56.6</td>
<td>41.2</td>
<td>23.1</td>
<td>42.0</td>
<td>50.3</td>
</tr>
<tr>



<td>BMGKD</td>
<td>39.4 (&#x002B;4.0)</td>
<td>57.3</td>
<td>41.9</td>
<td>23.7</td>
<td>42.8</td>
<td>51.0</td>
</tr>
<tr>
<td rowspan="15">RetinaNet</td>
<td rowspan="5">ResNet101</td>
<td rowspan="5">ResNet50</td>
<td>Baseline</td>
<td>37.4</td>
<td>55.4</td>
<td>39.1</td>
<td>20.4</td>
<td>40.3</td>
<td>48.1</td>
</tr>
<tr>



<td>FGD (2022)</td>
<td>39.1 (&#x002B;1.7)</td>
<td>59.0</td>
<td>42.3</td>
<td>22.8</td>
<td>43.1</td>
<td>52.3</td>
</tr>
<tr>



<td>CrossKD (2024)</td>
<td>39.7 (&#x002B;2.3)</td>
<td>58.9</td>
<td>42.5</td>
<td>22.4</td>
<td>43.6</td>
<td>52.8</td>
</tr>
<tr>



<td>HMKD (2026)</td>
<td>39.9 (&#x002B;2.5)</td>
<td>59.1</td>
<td>42.8</td>
<td>22.3</td>
<td>43.9</td>
<td>53.6</td>
</tr>
<tr>



<td>BMGKD</td>
<td>40.4 (&#x002B;2.8)</td>
<td>59.4</td>
<td>43.0</td>
<td>23.1</td>
<td>44.1</td>
<td>54.0</td>
</tr>
<tr>

<td rowspan="5">ResNet101</td>
<td rowspan="5">ResNet18</td>
<td>Baseline</td>
<td>31.6</td>
<td>49.6</td>
<td>33.3</td>
<td>16.1</td>
<td>34.2</td>
<td>43.0</td>
</tr>
<tr>



<td>FGD (2022)</td>
<td>33.9 (&#x002B;1.9)</td>
<td>51.6</td>
<td>35.2</td>
<td>18.3</td>
<td>36.3</td>
<td>44.9</td>
</tr>
<tr>



<td>CrossKD (2024)</td>
<td>34.1 (&#x002B;2.5)</td>
<td>52.2</td>
<td>36.1</td>
<td>18.6</td>
<td>36.9</td>
<td>47.4</td>
</tr>
<tr>



<td>HMKD (2026)</td>
<td>34.2 (&#x002B;2.6)</td>
<td>52.3</td>
<td>35.9</td>
<td>18.7</td>
<td>37.0</td>
<td>47.5</td>
</tr>
<tr>



<td>BMGKD</td>
<td>34.6 (&#x002B;3.0)</td>
<td>52.5</td>
<td>36.1</td>
<td>19.0</td>
<td>37.3</td>
<td>47.9</td>
</tr>
<tr>

<td rowspan="5">ResNet50</td>
<td rowspan="5">MobileNetV3</td>
<td>Baseline</td>
<td>31.4</td>
<td>49.3</td>
<td>33.1</td>
<td>15.8</td>
<td>34.1</td>
<td>42.7</td>
</tr>
<tr>



<td>MLD (2023)</td>
<td>34.1 (&#x002B;2.7)</td>
<td>52.0</td>
<td>35.9</td>
<td>18.4</td>
<td>37.0</td>
<td>46.9</td>
</tr>
<tr>



<td>PCD (2025)</td>
<td>34.2 (&#x002B;2.8)</td>
<td>52.1</td>
<td>36.0</td>
<td>18.6</td>
<td>37.1</td>
<td>46.9</td>
</tr>
<tr>



<td>CSAKD (2025)</td>
<td>34.5 (&#x002B;3.1)</td>
<td>52.4</td>
<td>36.3</td>
<td>18.8</td>
<td>37.2</td>
<td>47.1</td>
</tr>
<tr>



<td>BMGKD</td>
<td>34.8 (&#x002B;3.4)</td>
<td>52.8</td>
<td>36.5</td>
<td>19.4</td>
<td>37.6</td>
<td>47.6</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, on the Pascal VOC 2007 dataset, when using ResNet101 as the teacher network to train the student network ResNet50, BMGKD achieved the best performance across all three detectors. Faster R-CNN, RetinaNet, and GFL achieved mAPs of 56.2<inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, 58.0<inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and 58.3<inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, respectively. These results not only significantly outperformed the student baselines but also comprehensively surpassed the corresponding teacher models, maintaining a clear lead over existing methods such as AMD, BCKD, and SamKD. These results fully validate BMGKD&#x2019;s effectiveness across different datasets, its generalization capability, and its ability to elevate student models to a level surpassing that of the teacher.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance comparison on Pascal VOC 2007 dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Method</th>
<th colspan="3">Faster R-CNN</th>
<th colspan="3">RetinaNet</th>
<th colspan="3">GFL</th>
</tr>
<tr>

<th>mAP</th>
<th><inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th>mAP</th>
<th><inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th>mAP</th>
<th><inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>50</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msub><mml:mi>AP</mml:mi><mml:mrow><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>Teacher</td>
<td>56.3</td>
<td>82.7</td>
<td>62.6</td>
<td>58.2</td>
<td>82.0</td>
<td>63.0</td>
<td>58.4</td>
<td>81.6</td>
<td>64.3</td>
</tr>
<tr>
<td>Student</td>
<td>54.2</td>
<td>82.1</td>
<td>59.9</td>
<td>56.1</td>
<td>80.9</td>
<td>60.7</td>
<td>56.1</td>
<td>80.0</td>
<td>61.6</td>
</tr>
<tr>
<td>AMD [<xref ref-type="bibr" rid="ref-31">31</xref>] (2023)</td>
<td>55.3</td>
<td>82.3</td>
<td>61.4</td>
<td>57.2</td>
<td>81.2</td>
<td>62.1</td>
<td>57.5</td>
<td>80.6</td>
<td>62.9</td>
</tr>
<tr>
<td>BCKD [<xref ref-type="bibr" rid="ref-20">20</xref>] (2023)</td>
<td>55.6</td>
<td>82.4</td>
<td>61.9</td>
<td>57.5</td>
<td>81.4</td>
<td>62.5</td>
<td>57.9</td>
<td>80.9</td>
<td>63.3</td>
</tr>
<tr>
<td>SamKD [<xref ref-type="bibr" rid="ref-21">21</xref>] (2025)</td>
<td>55.9</td>
<td>82.5</td>
<td>62.3</td>
<td>57.8</td>
<td>81.6</td>
<td>62.9</td>
<td>58.1</td>
<td>81.1</td>
<td>63.7</td>
</tr>
<tr>
<td>BMGKD</td>
<td>56.2</td>
<td>82.6</td>
<td>62.4</td>
<td>58.0</td>
<td>81.8</td>
<td>63.1</td>
<td>58.3</td>
<td>81.3</td>
<td>64.0</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-3">Table 3</xref> compares the training overhead and accuracy gains of various distillation methods under the GFL ResNet101<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>ResNet18 setting. None of the methods alter the student architecture, and there are no additional parameters or computational costs during the inference stage. Compared to the response-based method CrossKD, the feature-based method FGD, and the hybrid method SamKD, BMGKD requires only 5.9 GFLOPs of additional computational effort and 1.14 times the training duration, achieving the highest mAP improvement of 4.6<inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and striking the optimal balance between training efficiency and accuracy.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Training cost and performance improvement of different KD methods.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Type</th>
<th>Additional GFLOPs (Training)</th>
<th>Training Time (Relative to Baseline)</th>
<th><inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> of mAP (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CrossKD [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Response-based</td>
<td>&#x002B;1.9</td>
<td>1.03<inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x002B;2.7%</td>
</tr>
<tr>
<td>FGD [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Feature-based</td>
<td>&#x002B;4.8</td>
<td>1.12<inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x002B;2.5%</td>
</tr>
<tr>
<td>SamKD [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>Response &#x002B; Feature</td>
<td>&#x002B;9.5</td>
<td>1.18<inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x002B;3.9%</td>
</tr>
<tr>
<td>BMGKD</td>
<td>Response &#x002B; Feature</td>
<td>&#x002B;5.9</td>
<td>1.14<inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x002B;4.6%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Hyperparameter Analysis</title>
<p>To determine the optimal loss weights, this section conducted hyperparameter sensitivity experiments using the control variable method based on the GFL detection framework, with ResNet-101 serving as the teacher model and ResNet-50 as the student model. As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, the optimal hyperparameter settings for BMGKD were determined to be <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.04</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.02</mml:mn></mml:math></inline-formula>, and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>. Under this configuration, the student model achieved a mAP of 44.4<inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on the GFL detection framework. Searches in different directions within the neighborhood of the optimal values did not reveal any performance gains. All deviations led to a consistent decline in detection accuracy.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Parameter sensitivity analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th><inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula></th>
<th>mAP (%)</th>
<th><inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula> from Best</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mn>44.4</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:mo>&#x2212;</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mn>0.75</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mn>0.75</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mn>44.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mo>&#x2212;</mml:mo><mml:mn>0.4</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mn>1.25</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mn>1.25</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mn>44.1</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mo>&#x2212;</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mn>1.5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mn>1.5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mn>44.2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:mo>&#x2212;</mml:mo><mml:mn>0.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mn>1.25</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mn>0.75</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:mn>44.2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mo>&#x2212;</mml:mo><mml:mn>0.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mn>0.02</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mn>0.03</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mn>44.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mo>&#x2212;</mml:mo><mml:mn>0.4</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mn>0.03</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mn>43.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mo>&#x2212;</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mn>0.06</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mn>0.03</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mn>43.8</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mo>&#x2212;</mml:mo><mml:mn>0.6</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mn>0.06</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mn>0.04</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mn>1.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mn>43.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mo>&#x2212;</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Ablation Experiments</title>
<p>To validate the effectiveness of each module in the BMGKD method, we used ResNet-101 and ResNet-50 as teacher-student backbone networks and conducted systematic ablation experiments based on the single-stage anchor-free detector GFL. The experimental results are shown in <xref ref-type="table" rid="table-5">Table 5</xref>. When using only the global feature difference alignment module, the model&#x2019;s detection accuracy improved by 2.8<inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. When using only the local feature difference alignment module, the model&#x2019;s detection accuracy improved by 2.6<inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. After combining the global and local feature difference modules, the accuracy improvement increased to 3.8<inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. Furthermore, when the feature difference module was integrated with the response output difference module, the model achieved optimal performance, with an mAP improvement of 4.2<inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> compared to the baseline model. The experimental results fully demonstrate that the complementary collaboration among the modules in the proposed method comprehensively improves the model&#x2019;s final detection accuracy.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Ablation experiments.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Global</th>
<th>Local</th>
<th>Response</th>
<th>mAP</th>
<th><inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>S</mml:mi></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>M</mml:mi></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>L</mml:mi></mml:msub></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>40.2</td>
<td>23.3</td>
<td>44.0</td>
<td>52.2</td>
</tr>
<tr>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>43.0 <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mo stretchy="false">(</mml:mo><mml:mo>+</mml:mo><mml:mn>2.8</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>26.0</td>
<td>46.5</td>
<td>55.1</td>
</tr>
<tr>
<td><inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>42.8 <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mo stretchy="false">(</mml:mo><mml:mo>+</mml:mo><mml:mn>2.6</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>26.2</td>
<td>46.5</td>
<td>54.9</td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>44.0 <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:mo stretchy="false">(</mml:mo><mml:mo>+</mml:mo><mml:mn>3.8</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>26.9</td>
<td>47.3</td>
<td>55.7</td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>44.4 <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mo stretchy="false">(</mml:mo><mml:mo>+</mml:mo><mml:mn>4.2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>27.4</td>
<td>49.0</td>
<td>56.8</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Data Visualization</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the object detection visualization results of the teacher model, student model, and the proposed BMGKD method across different scenarios. The pure student model exhibits significant performance shortcomings. It not only makes redundant errors such as misclassifying a bench as an object but also suffers from issues such as inflated confidence scores and insufficient bounding box accuracy, resulting in a significant gap in detection performance compared to the teacher model. In contrast, the BMGKD method, through its bidirectional mask-guided knowledge distillation mechanism, successfully transfers the teacher model&#x2019;s target discrimination capability to the student model. This not only eliminates false background detections in the student model but also calibrates the detection confidence to a level highly consistent with that of the teacher model, demonstrating greater robustness and detection reliability in scenarios with multiple targets and complex background interference.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Comparison of predicted results. (<bold>a</bold>) Original image. (<bold>b</bold>) Teacher model. (<bold>c</bold>) Student model. (<bold>d</bold>) BMGKD.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83923-fig-2.tif"/>
</fig>
<p>To visually demonstrate the proposed method&#x2019;s ability to focus on key regions of interest, we selected a representative test image from the MS COCO dataset and visualized the attention heatmap extracted from the GFL-ResNet50 student model after 12 epochs of training. The results are shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The intensity of the colors in the heatmap indicates the level of attention weighting, with red areas representing the highest level of attention and blue areas representing the lowest. As shown in the attention map for aircraft class detection, BMGKD can accurately focus on three key feature regions, the aircraft nose, engines, and tail, while filtering out redundant background interference. This fully validates that the collaborative masking design in the feature difference module can effectively enhance the student network&#x2019;s ability to understand and distinguish target features.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Attention comparison diagram. (<bold>a</bold>) Original image. (<bold>b</bold>) DMKD. (<bold>c</bold>) BMGKD.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83923-fig-3.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, we extracted feature maps from seven layers B3 and B4 in the backbone, and P2 through P6 in the neck of the GFL detection network. We calculated the channel averages layer by layer to obtain two-dimensional activation matrices, which were then visualized using a Cividis color map, where higher brightness indicates stronger activation. After BMGKD distillation, the feature activation distributions of the student model at each layer closely align with those of the teacher model. In the deeper layers P3 to P5, the student model successfully reproduces the semantic regions and object contours of interest to the teacher, whereas the activations of the original student model were relatively diffuse and lacked sufficient semantic distinctiveness. These results demonstrate that the global and local alignment mechanisms within the feature difference distillation module effectively transfer the teacher&#x2019;s tacit knowledge to the student model.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Feature visualization. (<bold>a</bold>) Original image. (<bold>b</bold>) Heatmap of feature channel activation values.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_83923-fig-4.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>To effectively improve the performance of knowledge distillation, we propose an object detection knowledge distillation algorithm designed to bridge the multidimensional differences between student and teacher networks. By simultaneously processing global and local features, this algorithm specifically addresses the differences in feature representations between the teacher network and the student network. In addition, at the model&#x2019;s response output layer, we designed a differentiated knowledge distillation strategy specific to the classification and regression tasks of object detection. This further bridges task-specific differences and enhances the student network&#x2019;s ability to understand the implicit knowledge of the teacher network.</p>
<p>Extensive comparative experiments, ablation experiments, and visualization results demonstrate that the proposed BMGKD algorithm delivers outstanding performance, with each module contributing positive gains, and exhibits strong competitiveness across various mainstream object detection frameworks. In future research, we will continue to advance lightweight feature distillation and loss function simplification to reduce resource consumption during training. At the same time, we will focus on exploring feature alignment and masking strategies that are more robust to extreme scales, occlusions, and long-tail classes, designing class-aware distillation mechanisms, and introducing learnable intermediate adaptation modules to further bridge the semantic gap across architectures.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Army Engineering University of PLA. The authors would like to acknowledge Army Engineering University of PLA for funding this work.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Tianqi Wang; data collection: Yang Li; analysis and interpretation of results: Tianqi Wang and Zhisong Pan; draft manuscript preparation: Zhisong Pan. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available in MS COCO at <ext-link ext-link-type="uri" xlink:href="https://cocodataset.org">https://cocodataset.org</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>P</given-names></string-name>, <string-name><surname>Jiao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A progressive semi-distillation model for dual-source remote sensing image classification</article-title>. <source>IEEE Trans Cybern</source>. <year>2026</year>;<volume>56</volume>(<issue>1</issue>):<fpage>67</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TCYB.2025.3616504</pub-id>; <pub-id pub-id-type="pmid">41091612</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Adaptive fine-grained fusion network for multimodal UAV object detection</article-title>. <source>IEEE Trans Image Process</source>. <year>2026</year>;<volume>35</volume>:<fpage>1870</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tip.2026.3661868</pub-id>; <pub-id pub-id-type="pmid">41671144</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hinton</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Dean</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Distilling the knowledge in a neural network</article-title>. <comment>arXiv:1503.02531. 2015</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Mimicking very efficient network for object detection</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>7341</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>G</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Han</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chandraker</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Learning efficient object detection models with knowledge distillation</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017</year>. p. <fpage>742</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Head-adaptive knowledge distillation for dense object detection</article-title>. In: <conf-name>2025 8th International Symposium on Big Data and Applied Statistics (ISBDAS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2025</year>. p. <fpage>857</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lan</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Gradient-guided knowledge distillation for object detectors</article-title>. In: <conf-name>2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>423</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Focal and global knowledge distillation for detectors</article-title>. In: <conf-name>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>4633</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Distilling object detectors with scale-conscious knowledge in remote sensing images</article-title>. <source>IEEE Geosci Remote Sens Lett</source>. <year>2026</year>;<volume>23</volume>:<fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/lgrs.2025.3634080</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Girshick</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Fast R-CNN</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>1440</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ricki</surname> <given-names>R</given-names></string-name>, <string-name><surname>Dicky</surname> <given-names>D</given-names></string-name>, <string-name><surname>Apriyansyah</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Fast region-based convolutional neural network in object detection: a review</article-title>. <source>Intl J Adv Comput Inform</source>. <year>2025</year>;<volume>2</volume>(<issue>1</issue>):<fpage>34</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.71129/ijaci.v2i1.pp34-40</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Focal loss for dense object detection</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>2980</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>F-YOLO: delving into fuzzy YOLO for improved traffic object detection</article-title>. <source>IEEE Trans Fuzzy Syst</source>. <year>2026</year>;<volume>34</volume>(<issue>2</issue>):<fpage>440</fpage>&#x2013;<lpage>52</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>He</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Fcos: fully convolutional one-stage object detection</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>9627</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Generalized focal loss V2: learning reliable localization quality estimation for dense object detection</article-title>. In: <conf-name>2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>11627</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Romero</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Fitnets: hints for thin deep nets</article-title>. <comment>arXiv:1412.6550. 2014</comment>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lv</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Wasserstein distance rivals kullback-leibler divergence for knowledge distillation</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2024</year>;<volume>37</volume>:<fpage>65445</fpage>&#x2013;<lpage>75</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Multi-scale feature fusion with knowledge distillation for object detection in aerial imagery</article-title>. <source>Eng Appl Artif Intell</source>. <year>2025</year>;<volume>158</volume>(<issue>2</issue>):<fpage>111518</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2025.111518</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Localization distillation for object detection</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2023</year>;<volume>45</volume>(<issue>8</issue>):<fpage>10070</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2023.3248583</pub-id>; <pub-id pub-id-type="pmid">37027640</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Bridging cross-task protocol inconsistency for distillation in dense object detection</article-title>. In: <conf-name>2023 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>17129</fpage>&#x2013;<lpage>38</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>SAMKD: spatial-aware adaptive masking knowledge distillation for object detection</article-title>. <comment>arXiv:2501.07101. 2025</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Saha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bialkowski</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khalifa</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Distilling representational similarity using centered kernel alignment (CKA)</article-title>. In: <conf-name>Proceedings of the the 33rd British Machine Vision Conference (BMVC 2022)</conf-name>. <publisher-loc>London, UK</publisher-loc>: <publisher-name>British Machine Vision Association</publisher-name>; <year>2022</year>. p. <fpage>1</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Maire</surname> <given-names>M</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hays</surname> <given-names>J</given-names></string-name>, <string-name><surname>Perona</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramanan</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Microsoft coco: common objects in context</article-title>. In: <conf-name>European Conference on Computer Vision</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2014</year>. p. <fpage>740</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Everingham</surname> <given-names>M</given-names></string-name>, <string-name><surname>Eslami</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Van Gool</surname> <given-names>L</given-names></string-name>, <string-name><surname>Williams</surname> <given-names>CK</given-names></string-name>, <string-name><surname>Winn</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>The pascal visual object classes challenge: a retrospective</article-title>. <source>Intl J Comput Vis</source>. <year>2015</year>;<volume>111</volume>(<issue>1</issue>):<fpage>98</fpage>&#x2013;<lpage>136</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-014-0733-5</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Task-balanced distillation for object detection</article-title>. <source>Pattern Recognit</source>. <year>2023</year>;<volume>137</volume>(<issue>2</issue>):<fpage>109320</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2023.109320</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lv</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name></person-group>. <article-title>DrKD: decoupling response-based distillation for object detection</article-title>. <source>Pattern Recognit</source>. <year>2025</year>;<volume>161</volume>:<fpage>111275</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2024.111275</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>CrossKD: cross-head knowledge distillation for object detection</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>16520</fpage>&#x2013;<lpage>30</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Remote sensing object detection through hierarchical feature mining and multivariate head collaboration with knowledge distillation</article-title>. <source>Neural Netw</source>. <year>2026</year>;<volume>195</volume>(<issue>10</issue>):<fpage>108205</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neunet.2025.108205</pub-id>; <pub-id pub-id-type="pmid">41110200</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Multi-level logit distillation</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>24276</fpage>&#x2013;<lpage>85</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Progressive class-level distillation</article-title>. <comment>arXiv:2505.24310. 2025</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Amd: adaptive masked distillation for object detection</article-title>. In: <conf-name>2023 International Joint Conference on Neural Networks (IJCNN)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>