<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">76623</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.076623</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>MobiIris: Attention-Enhanced Lightweight Iris Recognition with Knowledge Distillation and Quantization</article-title>
<alt-title alt-title-type="left-running-head">MobiIris: Attention-Enhanced Lightweight Iris Recognition with Knowledge Distillation and Quantization</alt-title>
<alt-title alt-title-type="right-running-head">MobiIris: Attention-Enhanced Lightweight Iris Recognition with Knowledge Distillation and Quantization</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Huynh</surname><given-names>Trong-Thua</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>thuaht@ptit.edu.vn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Huynh</surname><given-names>De-Thu</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Phu</surname><given-names>Du-Thang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Nguyen</surname><given-names>Hong-Son</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Nguyen</surname><given-names>Quoc H.</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Faculty of Information Technology II, Posts and Telecommunications Institute of Technology</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computer Science &#x0026; Engineering, The Saigon International University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
<aff id="aff-3"><label>3</label><institution>Institute of Digital Technology, Thu Dau Mot University</institution>, <addr-line>Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Trong-Thua Huynh. Email: <email>thuaht@ptit.edu.vn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>9</day><month>4</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>3</issue>
<elocation-id>19</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>11</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_76623.pdf"></self-uri>
<abstract>
<p>This paper introduces MobiIris, a lightweight deep network for mobile iris recognition that enhances attention and specifically addresses the balance between accuracy and efficiency on devices with limited resources. The proposed model is based on the large version of MobileNetV3 and adds more spatial attention blocks and an embedding-based head that was trained using margin-based triplet learning, enabling fine-grained modeling of iris textures in a compact representation. To further improve discriminability, we design a training pipeline that combines dynamic-margin triplet loss, a staged hard/semi-hard negative mining strategy, and feature-level knowledge distillation from a ResNet-50 teacher. Finally, we investigate the use of post-training float16 quantization to reduce memory footprint and latency for deployment on mobile hardware. Experiments on the challenging CASIA-IrisV4-Thousand dataset show that the full-precision MobiIris model requires only <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>12</mml:mn></mml:math></inline-formula> MB of storage and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mn>27</mml:mn></mml:math></inline-formula> ms inference latency, while achieving an EER of <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mn>1.409</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, VR@FAR &#x003D; <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>1</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mn>98.184</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and CMC@<inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mn>1</mml:mn></mml:math></inline-formula> of <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mn>94.785</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, closely matching a ResNet-50 baseline that is more than <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> larger and slower. Under post-training quantization, the model shrinks to <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mn>5.94</mml:mn></mml:math></inline-formula> MB with <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mn>13</mml:mn></mml:math></inline-formula> ms latency and maintains a competitive balance between accuracy and efficiency compared to other optimized variants. These results demonstrate that a coherent combination of lightweight architecture design, attention mechanisms, metric-learning objectives, hard negative mining, and knowledge distillation yields a practical iris recognition solution suitable for secure, real-time authentication on mobile and embedded platforms.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Iris recognition</kwd>
<kwd>lightweight architecture</kwd>
<kwd>model optimization</kwd>
<kwd>attention mechanism</kwd>
<kwd>knowledge distillation</kwd>
<kwd>model quantization</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Compared with other biometric features, the unique and stable nature of the iris pattern makes it a reliable and non-invasive means of identification [<xref ref-type="bibr" rid="ref-1">1</xref>]. This capability has proven effective across a wide range of security applications, such as mobile and physical access control, thereby contributing to the expansion of the global iris recognition market. Through miniaturized near-infrared sensors, iris recognition has been deployed on smartphones, which enables contactless authentication in response to the post-pandemic demand for touchless technologies [<xref ref-type="bibr" rid="ref-2">2</xref>]. Reflecting this momentum, Mordor Intelligence projects that this market will grow from USD <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mn>5.14</mml:mn></mml:math></inline-formula> billion in 2025 to USD <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mn>12.92</mml:mn></mml:math></inline-formula> billion by 2030, with a compound annual growth rate of <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mn>20.23</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-3">3</xref>]. Given this trend, accurate and efficient iris recognition solutions are increasingly important.</p>
<p>However, achieving a balance between accuracy and efficiency remains challenging. This is mainly because widely used state-of-the-art deep learning models are built on large convolutional architectures that demand substantial computational and memory resources. For instance, VGG-16 contains over <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mn>138</mml:mn></mml:math></inline-formula> million parameters and requires several gigabytes of memory during inference [<xref ref-type="bibr" rid="ref-4">4</xref>]. Moreover, DenseNet-121, compared with most ResNet architectures, is more parameter-efficient, with about <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mn>8</mml:mn></mml:math></inline-formula> million parameters, yet still has a memory footprint of hundreds of megabytes [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. This high computational requirement makes mobile deployment difficult for iris recognition, as small and fast models are critical for user experience and scalability. Moreover, iris acquisition in uncontrolled conditions poses challenges, such as illumination variation, occlusion, and reflection, which can significantly degrade recognition performance.</p>
<p>Over the years, iris recognition research has steadily progressed, moving from classical image processing and handcrafted machine learning features to deep learning methods [<xref ref-type="bibr" rid="ref-7">7</xref>]. Despite these improvements, the high computational demand of traditional models hinders their deployment on resource-constrained devices, thereby motivating growing interest in lightweight architectures. On the ImageNet dataset, GhostNet, a lightweight architecture with an improved feature generation module, achieved the highest classification accuracy among lightweight models, with competitive computational cost [<xref ref-type="bibr" rid="ref-8">8</xref>]. Similarly, [<xref ref-type="bibr" rid="ref-9">9</xref>] evaluated a set of deep learning models for iris recognition under constrained conditions, showing that MobileNetV3 models reduced parameters by <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mn>87</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>/<inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>83</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and shortened inference time by <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mn>80</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>/<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>56</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> compared to ResNet-50/DenseNet-201, respectively, on OpenEDS dataset. Beyond the native lightweight architectures, deep convolutional networks have also been adapted into lite versions to achieve comparable performance with lower resource usage. For example, ResNet-Lite applies model optimization techniques to compress ResNet-50 while maintaining near-original performance, achieving improvements of <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>5.40</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mn>7.13</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on CIFAR-10 and Fashion-MNIST datasets, respectively [<xref ref-type="bibr" rid="ref-10">10</xref>]. In general, these findings show that lightweight architectures can deliver high accuracy with reduced computational cost.</p>
<p>While lightweight architectures establish a solid foundation for mobile iris recognition, further enhancements in performance can be obtained through optimization techniques. Introduced in [<xref ref-type="bibr" rid="ref-11">11</xref>], knowledge distillation involves transferring knowledge from a larger and more complex teacher model to a smaller and simpler student model. Despite being widely applied in various computer vision tasks, it is less explored in iris recognition. In [<xref ref-type="bibr" rid="ref-12">12</xref>], a comprehensive survey covering a wide range of models and knowledge distillation techniques reported accuracy improvements of over <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mn>5</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on CIFAR-10 and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mn>10</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on CIFAR-100, and [<xref ref-type="bibr" rid="ref-13">13</xref>] underscored the effectiveness of this technique in resource-constrained scenarios. Based on these findings, [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed a multimodal compression approach that incorporates knowledge distillation with additional optimization techniques. This study utilized a teacher model to guide the training of a VGG16 student model, thereby increasing accuracy by more than <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mn>1</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> (from <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mn>90</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mn>91</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>) while reducing model size from <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mn>528.95</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mn>90.22</mml:mn></mml:math></inline-formula> MB and inference time from <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mn>0.92</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mn>0.58</mml:mn></mml:math></inline-formula> s on CASIA-IrisV1 dataset. Generally, these results show that knowledge distillation enables smaller models to achieve performance comparable to larger models, which opens opportunities for further advances.</p>
<p>Despite recent notable progress, several important gaps persist in iris recognition research. First, although lightweight architectures and compression techniques have been widely adopted in other biometric modalities, such as face and fingerprint recognition, their use in iris recognition remains limited. Few studies have investigated their potential, thereby hindering the deployment of compact models in resource-constrained conditions. Second, while knowledge distillation has shown effectiveness in various computer vision tasks, its application combined with compression techniques in iris recognition remains relatively unexplored, limiting the development of compact and accurate models. Third, embedding-based feature learning, which is a common approach in most biometric modalities, has not been fully utilized in iris recognition, as most studies rely on classification-based outputs. Finally, balancing accuracy and efficiency remains challenging, as most prior works focus on one aspect, and few address both within a unified framework.</p>
<p>Motivated by these challenges, this study aims to balance accuracy and efficiency for practical deployment. Specifically, the main objectives are: <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to propose a lightweight deep learning model optimized for mobile iris recognition; <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo stretchy="false">(</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to improve model performance through optimization methods; <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to provide a comprehensive evaluation of model performance on a standard dataset. Accordingly, the primary contributions are: <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> a MobileNetV3-based model enhanced by spatial attention mechanism and task-specific embedding output module, featuring a model size under <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mn>10</mml:mn></mml:math></inline-formula> MB and inference latency below <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mn>50</mml:mn></mml:math></inline-formula> ms; <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mo stretchy="false">(</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> task-specific adaptation of hard negative mining strategy and knowledge distillation technique to improve model performance; <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> competitive performance against the ResNet-50 model on the CASIA-IrisV4-Thousand dataset, with an Equal Error Rate of no more than <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mn>5</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a Verification Rate of at least <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mn>95</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> at <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mn>1</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> False Acceptance Rate, and a rank-<inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mn>1</mml:mn></mml:math></inline-formula> Cumulative Matching Characteristic of at least <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mn>90</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methodology</title>
<sec id="s2_1">
<label>2.1</label>
<title>Dataset Preparation</title>
<p>For iris recognition tasks, datasets should be carefully prepared to ensure both diversity and consistency. This study employs the CASIA-IrisV4 dataset [<xref ref-type="bibr" rid="ref-15">15</xref>], which contains near-infrared images collected under various acquisition conditions. Specifically, the original data are reorganized into a subject-based directory structure in which each image folder belongs to a single subject. Since the focus of this work is on iris recognition, including feature extraction, encoding, and matching stages, iris segmentation is omitted. Instead, we propose a preprocessing procedure (Algorithm 1) that utilizes Open-Iris library to automatically detect and extract the iris region from each input image for recognition purposes, thereby ensuring that all modules operate solely on the normalized iris area [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>For our experiments, we conducted tests on all CASIA-IrisV4 subsets. CASIA-IrisV4-Thousand was selected because it yielded the most representative recognition results. This subset also presents challenging acquisition conditions and is widely used in previous iris recognition benchmarks. In detail, it consists of 1000 subjects with a total of 20,000 images, about <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mn>20</mml:mn></mml:math></inline-formula> images per subject. After processing with our proposed procedure (Algorithm 1), the final dataset contains <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mn>1997</mml:mn></mml:math></inline-formula> identities, corresponding to the split of left and right eyes into separate identities, and 18,800 images in total, reflecting the removal of improperly processed images. Approximately <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mn>1200</mml:mn></mml:math></inline-formula> images, accounting for roughly <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mn>6</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of the dataset, were discarded due to segmentation failures, mainly caused by eyelid or eyelash occlusion, low iris&#x2013;boundary contrast, or inaccurate pupil localization produced by Open-Iris. On average, each identity includes about <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mn>10</mml:mn></mml:math></inline-formula> images ranging from <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mn>2</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mn>10</mml:mn></mml:math></inline-formula>. Example preprocessed iris region images are shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Iris images before and after preprocessing.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-1.tif"/>
</fig>
<fig id="fig-6">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-6.tif"/>
</fig>
<p>Algorithm 1 presents the iris-image preprocessing pipeline, providing key operations, starting from loading the input image, extracting the relevant regions to create the region of interest mask, computing the iris bounding box to cropping the iris region, to generating the final output. To ensure uniformity across all samples, additional steps are included to discard improperly processed images and separate left and right eyes of each subject into independent identities, which play a crucial role in model training.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Model Architecture</title>
<p>To balance accuracy and efficiency, the large variant of MobileNetV3 is adopted as the backbone for its sufficient representational capacity. Architecturally, an embedding-based module generates compact feature embeddings, while Spatial Attention (SA) complements Squeeze-and-Excitation (SE) to enhance spatial features. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, SA provides a more efficient alternative to Convolutional Block Attention Module (CBAM) and Non-Local Attention (NLA).</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of attention variants by functionality and complexity on MobileNetV3-L.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th></th>
<th><inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>F</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>P</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>s</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>F</mml:mi><mml:mi>L</mml:mi><mml:mi>O</mml:mi><mml:mi>P</mml:mi><mml:mi>S</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mn>6</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>SA</italic></td>
<td><italic>Local Spatial</italic></td>
<td><inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mn>0.049</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mn>2.4</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>38.4</mml:mn></mml:math></inline-formula></td>
<td><italic>Lightweight and Sufficient</italic></td>
</tr>
<tr>
<td><italic>CBAM</italic></td>
<td><italic>Channel &#x0026; Spatial</italic></td>
<td><inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mn>0.072</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>3.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mn>5.6</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>38.5</mml:mn></mml:math></inline-formula></td>
<td><italic>Partially Redundant</italic></td>
</tr>
<tr>
<td><italic>NLA</italic></td>
<td><italic>Global Spatial</italic></td>
<td><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>1.73</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>76.8</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mn>384</mml:mn><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>14</mml:mn><mml:mo>,</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>752</mml:mn></mml:math></inline-formula></td>
<td><italic>High Computational Cost</italic></td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Backbone Module</title>
<p>The Spatial Attention, illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, functions similarly to the SE Attention; however, it is designed to focus on spatial features. Given an intermediate feature map <italic>F</italic>, spatial information is aggregated by applying average pooling and max pooling across the channel dimension, producing two spatial context descriptors, <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>.<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">(</mml:mo></mml:mrow></mml:mstyle><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Spatial attention module for discriminative region highlighting.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-2.tif"/>
</fig>
<p>These context descriptors are then combined to generate the spatial attention map <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>M</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>F</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula><inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> using the sigmoid function <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></inline-formula>. Here, <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msup><mml:mi>A</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents a <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi></mml:math></inline-formula> convolution kernel used to aggregate local spatial information, with smaller <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>n</mml:mi></mml:math></inline-formula> capturing local details and larger <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>n</mml:mi></mml:math></inline-formula> incorporating broader spatial context.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Head Module</title>
<p>Instead of relying on closed-set classification, an embedding-based approach suited to the open-set nature of biometric recognition is adopted. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, its design consists of a linear layer followed by batch normalization, producing compact feature embeddings. In related biometric tasks, embedding-based methods outperform classification-based ones: ArcFace achieved state-of-the-art performance, reporting <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mn>99.83</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> accuracy on LFW and <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mn>96.98</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> VR@FAR of <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> on MegaFace [<xref ref-type="bibr" rid="ref-17">17</xref>]; SphereFace improved verification from <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mn>98.71</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> using softmax to <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mn>99.42</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on LFW [<xref ref-type="bibr" rid="ref-18">18</xref>]; and CosFace obtained an accuracy of <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mn>99.73</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on LFW and <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mn>97.6</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on YTF [<xref ref-type="bibr" rid="ref-19">19</xref>]. These methods train deep networks with margin-based losses, such as angular margin in SphereFace, cosine margin in CosFace, and additive angular margin in ArcFace, to project features into a hyperspherical space that enhances inter-class separability and intra-class compactness.<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mi>v</mml:mi><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:msqrt><mml:msubsup><mml:mi>v</mml:mi><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mi>n</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:msqrt></mml:mfrac></mml:math></disp-formula></p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Embedding-based module for feature representation learning.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-3.tif"/>
</fig>
<p>Let <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>v</mml:mi></mml:math></inline-formula> be an embedding vector of dimension <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>n</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> normalization <xref ref-type="disp-formula" rid="eqn-2">(2)</xref> is applied by dividing each element of <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>v</mml:mi></mml:math></inline-formula> by its norm. After normalization, the resulting vectors satisfy unit length, <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>v</mml:mi><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, giving all embeddings the same magnitude. This property ensures that the model focuses on the direction of the embeddings rather than their scale, since the direction encodes the identity-related information while the magnitude mostly reflects feature scale.</p>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Model Optimization</title>
<p>Together, the lightweight backbone, spatial attention module, and embedding-based head form a cohesive feature extraction and representation framework, which is further reinforced during training through Hard Negative Mining and Knowledge Distillation, enabling the model to focus on challenging samples and to transfer discriminative knowledge from a stronger teacher model.</p>
<p>Algorithm 2 outlines the overall training optimization procedure under the triplet-learning paradigm, including triplet batch sampling and processing. The sampling stage constructs a triplet consisting of an anchor (the reference sample), a positive sample with the same label, and a negative sample with a different label that has not been paired with the selected anchor and positive samples in the current iteration. The processing stage extracts student and teacher embeddings and applies hard negative mining with knowledge distillation, depending on the current phase.</p>
<fig id="fig-7">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-7.tif"/>
</fig>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>Triplet-Based Learning</title>
<p>Triplet-based learning is employed to learn discriminative embeddings by enforcing a relative distance constraint among anchor, positive and negative samples. Specifically, the model is trained to minimize the Euclidean distance <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>d</mml:mi></mml:math></inline-formula> between the anchor <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mi>a</mml:mi></mml:math></inline-formula> and positive <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mi>p</mml:mi></mml:math></inline-formula> while maximizing the distance between the anchor <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi>a</mml:mi></mml:math></inline-formula> and negative <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>n</mml:mi></mml:math></inline-formula> at least a margin <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mi>m</mml:mi></mml:math></inline-formula>, as formalized in the triplet margin loss <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-3">(3)</xref>. This encourages the embedding space to preserve identity-level similarity, ensuring that samples of the same class are closer together than samples of different classes.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>0</mml:mn><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>To optimize the triplet margin loss, a margin scheduling strategy <xref ref-type="disp-formula" rid="eqn-5">(5)</xref> is adopted, where the margin dynamically changes during training. The margin is initially fixed during a warm-hard phase and then gradually transitions to semi-hard phases using a cosine annealing schedule <xref ref-type="disp-formula" rid="eqn-4">(4)</xref>, facilitating effective integration with the hard negative mining strategy. Let <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:math></inline-formula> at epoch <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:mi>t</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the fixed warm-hard margin, and <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mi>M</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> are the start and end margins for the first and second halves of the semi-hard phase, respectively.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>A</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>A</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>A</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Hard Negative Mining</title>
<p>To enhance convergence and embedding discriminability, hard negative mining strategy <xref ref-type="disp-formula" rid="eqn-6">(6)</xref> is applied during training. In the initial warm-up stage, <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mi>W</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>H</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>M</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi></mml:math></inline-formula> is used, where semi-hard negatives that beyond the positive but within slightly loosened boundaries <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to avoid overly difficult cases are selected. In the remaining stages of the training, <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>v</mml:mi><mml:mi>e</mml:mi><mml:mi>M</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi></mml:math></inline-formula> is used, where both <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are gradually tightened to focus on negatives closer to the anchor, thereby progressively increasing the difficulty of the sampled pairs. When no samples satisfy the warm-hard or semi-hard condition, the top-k hardest negatives are used as substitutes. This progressive strategy enables the model to gradually transition from easier to more challenging samples.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi>W</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mo>:</mml:mo><mml:mspace width="1em" /><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mo>:</mml:mo><mml:mspace width="1em" /><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mo>:</mml:mo><mml:mspace width="1em" /><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2264;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s2_3_3">
<label>2.3.3</label>
<title>Knowledge Distillation</title>
<p>To strengthen the discriminative capability of the student model, MobileNetV3-L, knowledge distillation is employed using a pretrained teacher model, ResNet-50, trained on the same dataset, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The student model is optimized with a combined objective <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-7">(7)</xref>, including a triplet loss <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-3">(3)</xref> and a cosine similarity-based loss <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-8">(8)</xref> scaled by a factor <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> for each embedding type <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>x</mml:mi></mml:math></inline-formula>, and an overall factor <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-9">(9)</xref>.<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mi>&#x03B1;</mml:mi><mml:mn>3</mml:mn></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>n</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mi>x</mml:mi></mml:msub><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mi>x</mml:mi></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mi>x</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>y</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>x</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msubsup><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>x</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Knowledge Distillation between ResNet-50 teacher and MobileNetV3-L student.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-4.tif"/>
</fig>
<p>Let <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the dynamic weighting factor at epoch <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>t</mml:mi></mml:math></inline-formula>, with <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> for controlling the contribution of teacher model during the semi-hard phase. During the initial warm-up phase, only <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is applied to allow the student model to first learn basic features without being influenced by the teacher model. Then, each <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mi>x</mml:mi></mml:msubsup></mml:math></inline-formula> is added to progressively align the student representations with the teacher.
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>A</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Evaluation Metrics</title>
<p>To evaluate recognition performance, two sets of metrics that capture different aspects of the model capabilities are used. These metrics offer a comprehensive assessment of how accurately the model can verify or identify subjects under various conditions. By integrating both threshold-based and ranking-based measures, the evaluation quantifies not only the accuracy but also the robustness of the model in challenging scenarios, such as hard negatives and unseen identities.</p>
<sec id="s2_4_1">
<label>2.4.1</label>
<title>Similarity Metrics</title>
<p>A key step in recognition tasks is measuring the similarity between two normalized embedding vectors, <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, using either Cosine similarity <xref ref-type="disp-formula" rid="eqn-10">(10)</xref> or Euclidean distance <xref ref-type="disp-formula" rid="eqn-11">(11)</xref>. The similarity scores, <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, are used to compute threshold-based metrics such as EER and VR@FAR, for verification, and ranking-based metrics such as CMC@K, for identification.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mstyle displaystyle="true" scriptlevel="0"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mstyle><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mspace width="thinmathspace" /><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>b</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s2_4_2">
<label>2.4.2</label>
<title>Verification Metrics</title>
<p>Verification metrics are used to evaluate the model&#x2019;s ability to determine whether pairs of samples correspond to the same or different identities. Given a similarity score <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> or distance <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, a threshold <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> is applied to classify a pair as a match or non-match.
<list list-type="bullet">
<list-item>
<p>False Acceptance Rate, defined as <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mi>F</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></inline-formula>, measures the proportion of imposter pairs that are incorrectly accepted as matches, where <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the number of imposter pairs whose similarity exceeds the threshold <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> (or distance is below <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>), and <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the total number of imposter pairs.</p></list-item>
<list-item>
<p>False Rejection Rate, defined as <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mi>F</mml:mi><mml:mi>R</mml:mi><mml:mi>R</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></inline-formula>, measures the proportion of genuine pairs that are incorrectly rejected as non-matches, where <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the number of genuine pairs whose similarity is below the threshold <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> (or distance is above <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>), and <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the total number of genuine pairs.</p></list-item>
<list-item>
<p>Equal Error Rate, displayed as <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>F</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mi>R</mml:mi><mml:mi>R</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, represents the threshold <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> where false acceptance rate FAR and false rejection rate FRR are equal, which reflects the difference between incorrectly accepting imposter pairs and rejecting genuine pairs.</p></list-item>
<list-item>
<p>Verification Rate at a False Acceptance Rate, defined as <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mi>V</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></inline-formula>, measures the proportion of genuine pairs that model correctly verified at a defined FAR, where <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the number of genuine pairs accepted at <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> corresponding to FAR and <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the total number of genuine pairs.</p></list-item>
</list></p>
</sec>
<sec id="s2_4_3">
<label>2.4.3</label>
<title>Identification Metrics</title>
<p>Identification metrics are used to evaluate the model&#x2019;s ability to correctly identify a probe iris image from a gallery of known identities. In contrast to verification, which makes a binary decision, identification ranks all gallery images according to their similarity to the probe.
<list list-type="bullet">
<list-item>
<p>Cumulative Matching Characteristic, defined as <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mi>C</mml:mi><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:math></inline-formula>, measures the proportion that the correct identity of a probe appears is retrieved within the top-k ranked gallery results, where <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the rank threshold, <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the total number of probe images, and <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the number of probes whose correct match appears in the top-k gallery images.</p></list-item>
</list></p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Results and Discussion</title>
<sec id="s3_1">
<label>3.1</label>
<title>Experimental Setup</title>
<p>Experiments are conducted on the CASIA-IrisV4-Thousand dataset, split into <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mn>3</mml:mn></mml:math></inline-formula> subsets with a ratio of <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mn>70</mml:mn></mml:math></inline-formula>:<inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mn>15</mml:mn></mml:math></inline-formula>:<inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:mn>15</mml:mn></mml:math></inline-formula> at subject level to ensure that all images of the same subject are grouped together. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, the splits contain <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>397</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mn>300</mml:mn></mml:math></inline-formula>, and <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mn>300</mml:mn></mml:math></inline-formula> subjects, respectively, corresponding to 13,150, <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:mn>2838</mml:mn></mml:math></inline-formula>, and <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mn>2812</mml:mn></mml:math></inline-formula> images. Afterwards, all grayscale images were resized to <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mn>224</mml:mn></mml:math></inline-formula><inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula><inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mn>224</mml:mn></mml:math></inline-formula> and replicated across three channels for backbone compatibility, then scaled to the <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> range, finally normalized with the standard ImageNet mean and standard deviation. For data generalization, minor augmentation randomly applied rotations of within <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mo>&#x00B1;</mml:mo><mml:msup><mml:mn>5</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and affine transformations with translation and scaling within <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:mo>&#x00B1;</mml:mo><mml:mn>5</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> to avoid excessive distortion.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Overview of the Preprocessed Dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Subject</th>
<th>Image</th>
<th>Min Imgs/Sub</th>
<th>Max Imgs/Sub</th>
</tr>
</thead>
<tbody>
<tr>
<td>Training</td>
<td><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:mn>1397</mml:mn></mml:math></inline-formula></td>
<td>13,150</td>
<td><inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:mn>2</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:mn>10</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Validation</td>
<td><inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mn>300</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mn>2838</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mn>10</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Testing</td>
<td><inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mn>300</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mn>2812</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:mn>3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:mn>10</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td>Total</td>
<td><inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:mn>1997</mml:mn></mml:math></inline-formula></td>
<td>18,800</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The training process is conducted over <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:mn>40</mml:mn></mml:math></inline-formula> epochs with a batch size of <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:mn>32</mml:mn></mml:math></inline-formula>, divided into 3 stages: a warm-hard negative stage for the first <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:mn>8</mml:mn></mml:math></inline-formula> epochs, and a semi-hard negative stage divided evenly into two parts covering the remaining <inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:mn>32</mml:mn></mml:math></inline-formula> epochs. At the core of the training pipeline, the model is trained to minimize the <inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:math></inline-formula> <xref ref-type="disp-formula" rid="eqn-3">(3)</xref> with a dynamic margin varying between <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:mn>0.4</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:mn>0.5</mml:mn></mml:math></inline-formula>. Additionally, the optimization is performed using <inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>W</mml:mi></mml:math></inline-formula> with different learning rates <inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> and weight decays <inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula>: a lower learning rate is applied to the backbone <inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:mo>:</mml:mo><mml:mn>5</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>&#x03BB;</mml:mi><mml:mo>:</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to preserve pretrained features, while higher learning rates are used to the attention <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:mo>:</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>&#x03BB;</mml:mi><mml:mo>:</mml:mo><mml:mn>5</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and head <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B7;</mml:mi><mml:mo>:</mml:mo><mml:mn>5</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>5</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>&#x03BB;</mml:mi><mml:mo>:</mml:mo><mml:mn>5</mml:mn><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to learn task-specific representations. To leverage the pretrained backbone, its layers are gradually unfrozen for smooth fine-tuning, and two learning rate schedulers <xref ref-type="disp-formula" rid="eqn-12">(12)</xref> are applied: a short linear decay during warm-up phase to stabilize early training, followed by a cosine decay to help the model converge smoothly in the remaining stages.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>&#x03B7;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mfrac><mml:mi>t</mml:mi><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>,</mml:mo></mml:mstyle></mml:mtd><mml:mtd><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03C0;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mstyle></mml:mtd><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>w</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2264;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Regarding the hard negative mining strategy, the configuration defines the lower and upper bounds for negative sample selection as scales of the triplet margin. Specifically, these bounds are computed as the product of the margin and generated ratios, as formalized in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>, thereby selecting negatives according to the current discriminative ability of the model.
<disp-formula id="ueqn-13"><mml:math id="mml-ueqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the warm-hard stage, the ratios for the lower and upper bounds, <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mtext>left</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mtext>right</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, are fixed at <inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:mn>0.2</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:mn>1.0</mml:mn></mml:math></inline-formula>, respectively. During the semi-hard stage, these ratios are adjusted using a cosine decay scheduler from <inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:mn>0.2</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:mn>0.1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:mn>1.0</mml:mn></mml:math></inline-formula> to <inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:mn>0.9</mml:mn></mml:math></inline-formula>. Meanwhile, the topk-hardest strategy is applied to select one of the <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:mn>5</mml:mn></mml:math></inline-formula> hardest samples if no negatives satisfy the bounds.</p>
<p>For the knowledge distillation technique, the configuration specifies a maximum distillation weight <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> of <inline-formula id="ieqn-212"><mml:math id="mml-ieqn-212"><mml:mn>0.5</mml:mn></mml:math></inline-formula>, with all component weights <inline-formula id="ieqn-213"><mml:math id="mml-ieqn-213"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> set to <inline-formula id="ieqn-214"><mml:math id="mml-ieqn-214"><mml:mn>1.0</mml:mn></mml:math></inline-formula>, as defined in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>. This setup enables the student model to effectively leverage the teacher&#x2019;s guidance across the anchor, positive, and negative components, facilitating the learning of more discriminative embeddings and improving overall representation quality during training.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Experimental Result</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Original Performance</title>
<p>To evaluate the impact of the proposed optimizations, the baseline MobileNetV3-L (BASE) is compared with its optimized variants, and the state-of-the-art ResNet-50 (SOTA). Specifically, attention mechanism and knowledge distillation technique are incorporated into the baseline to create more robust variants. As observed, the individual variants show moderate improvements, whereas their combination (our proposal) yields a more substantial improvement and narrows the performance gap to the state-of-the-art. <xref ref-type="table" rid="table-3">Table 3</xref> presents a detailed comparison among the baseline, the optimized variants, and the state-of-the-art model.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance comparison across different optimization techniques at FAR &#x003D; 1% and K &#x003D; 1.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th></th>
<th><inline-formula id="ieqn-215"><mml:math id="mml-ieqn-215"><mml:mi>E</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-216"><mml:math id="mml-ieqn-216"><mml:mi>V</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-217"><mml:math id="mml-ieqn-217"><mml:mi>C</mml:mi><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>K</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-218"><mml:math id="mml-ieqn-218"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>M</mml:mi><mml:mi>B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-219"><mml:math id="mml-ieqn-219"><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>MobileNetV3-L (BASE)</italic></td>
<td><inline-formula id="ieqn-220"><mml:math id="mml-ieqn-220"><mml:mn>1.699</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-221"><mml:math id="mml-ieqn-221"><mml:mn>97.678</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-222"><mml:math id="mml-ieqn-222"><mml:mn>93.073</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-223"><mml:math id="mml-ieqn-223"><mml:mn>11.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-224"><mml:math id="mml-ieqn-224"><mml:mn>25.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Only Attention</italic></td>
<td><inline-formula id="ieqn-225"><mml:math id="mml-ieqn-225"><mml:mn>1.484</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-226"><mml:math id="mml-ieqn-226"><mml:mn>97.993</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-227"><mml:math id="mml-ieqn-227"><mml:mn>93.949</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-228"><mml:math id="mml-ieqn-228"><mml:mn>12.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-229"><mml:math id="mml-ieqn-229"><mml:mn>27.1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Only Distilled</italic></td>
<td><inline-formula id="ieqn-230"><mml:math id="mml-ieqn-230"><mml:mn>1.535</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-231"><mml:math id="mml-ieqn-231"><mml:mn>97.951</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-232"><mml:math id="mml-ieqn-232"><mml:mn>94.466</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-233"><mml:math id="mml-ieqn-233"><mml:mn>11.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-234"><mml:math id="mml-ieqn-234"><mml:mn>25.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Attention &#x0026; Distilled</italic></td>
<td><inline-formula id="ieqn-235"><mml:math id="mml-ieqn-235"><mml:mn>1.409</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-236"><mml:math id="mml-ieqn-236"><mml:mn>98.184</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-237"><mml:math id="mml-ieqn-237"><mml:mn>94.785</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-238"><mml:math id="mml-ieqn-238"><mml:mn>12.0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-239"><mml:math id="mml-ieqn-239"><mml:mn>27.1</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>ResNet-50 (SOTA)</italic></td>
<td><inline-formula id="ieqn-240"><mml:math id="mml-ieqn-240"><mml:mn>1.567</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-241"><mml:math id="mml-ieqn-241"><mml:mn>98.092</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-242"><mml:math id="mml-ieqn-242"><mml:mn>95.103</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-243"><mml:math id="mml-ieqn-243"><mml:mn>90.9</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-244"><mml:math id="mml-ieqn-244"><mml:mn>188.1</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The baseline achieves an EER of <inline-formula id="ieqn-245"><mml:math id="mml-ieqn-245"><mml:mn>1.699</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR of <inline-formula id="ieqn-246"><mml:math id="mml-ieqn-246"><mml:mn>97.678</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K of <inline-formula id="ieqn-247"><mml:math id="mml-ieqn-247"><mml:mn>93.073</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, while remaining highly efficient at only <inline-formula id="ieqn-248"><mml:math id="mml-ieqn-248"><mml:mn>11.9</mml:mn></mml:math></inline-formula> MB and <inline-formula id="ieqn-249"><mml:math id="mml-ieqn-249"><mml:mn>25.2</mml:mn></mml:math></inline-formula> ms latency. In contrast, the state-of-the-art outperforms the baseline in accuracy, achieving an EER of <inline-formula id="ieqn-250"><mml:math id="mml-ieqn-250"><mml:mn>1.567</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR of <inline-formula id="ieqn-251"><mml:math id="mml-ieqn-251"><mml:mn>98.092</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K of <inline-formula id="ieqn-252"><mml:math id="mml-ieqn-252"><mml:mn>95.103</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, at the expense of substantially larger size of <inline-formula id="ieqn-253"><mml:math id="mml-ieqn-253"><mml:mn>90.9</mml:mn></mml:math></inline-formula> MB and higher latency of <inline-formula id="ieqn-254"><mml:math id="mml-ieqn-254"><mml:mn>188.1</mml:mn></mml:math></inline-formula> ms. This comparison reflects the inherent trade-off between recognition accuracy and computational efficiency. To overcome this limitation, attention mechanism and knowledge distillation technique are applied, with the goal of pushing its accuracy closer to the state-of-the-art performance while retaining its compact-size and low-latency characteristics.</p>
<p>Compared to the baseline, both attention and distillation enhancements achieve consistent accuracy improvements with minimal computational cost. The attention-only variant reduces the EER to <inline-formula id="ieqn-255"><mml:math id="mml-ieqn-255"><mml:mn>1.484</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and increases both the VR@FAR and CMC@K to <inline-formula id="ieqn-256"><mml:math id="mml-ieqn-256"><mml:mn>97.993</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-257"><mml:math id="mml-ieqn-257"><mml:mn>93.949</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, respectively, as the attention mechanism helps the model capture more discriminative features. Similarly, the distilled-only variant achieves an EER of <inline-formula id="ieqn-258"><mml:math id="mml-ieqn-258"><mml:mn>1.535</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR of <inline-formula id="ieqn-259"><mml:math id="mml-ieqn-259"><mml:mn>97.951</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K of <inline-formula id="ieqn-260"><mml:math id="mml-ieqn-260"><mml:mn>94.466</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, showing the benefit of knowledge transfer in enhancing feature discrimination. Taken together, these results confirm that each enhancement, regardless of the approach, independently strengthens the model&#x2019;s discriminative ability.</p>
<p>Based on these individual enhancements, combining attention and distillation further improves overall performance. Compared to the baseline, the combined variant slightly increases model size by <inline-formula id="ieqn-261"><mml:math id="mml-ieqn-261"><mml:mn>0.1</mml:mn></mml:math></inline-formula> MB and latency by <inline-formula id="ieqn-262"><mml:math id="mml-ieqn-262"><mml:mn>2</mml:mn></mml:math></inline-formula> ms, while achieving superior performance across all metrics, with an EER of <inline-formula id="ieqn-263"><mml:math id="mml-ieqn-263"><mml:mn>1.409</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR of <inline-formula id="ieqn-264"><mml:math id="mml-ieqn-264"><mml:mn>98.184</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K of <inline-formula id="ieqn-265"><mml:math id="mml-ieqn-265"><mml:mn>94.785</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>. Moreover, it remains over <inline-formula id="ieqn-266"><mml:math id="mml-ieqn-266"><mml:mn>7</mml:mn><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> smaller and faster than the state-of-the-art model, achieving comparable recognition accuracy. Additionally, the combined model exhibits clear convergence during training, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. These results highlight the complementary advantages of these two techniques, since the attention strengthens feature discrimination and the distillation fosters generalizable representations. Collectively, this combination provides an excellent trade-off between accuracy and efficiency, making the model highly practical for mobile iris recognition.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison of training and validation convergence.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_76623-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Quantized Performance</title>
<p>To further examine the practicality of the proposed models for mobile deployment, we assess their performance under post-training quantization. As summarized in <xref ref-type="table" rid="table-4">Table 4</xref>, the results highlight that reduced numerical precision substantially impacts recognition accuracy and computational efficiency. As expected, quantization causes a noticeable degradation in recognition performance across all models.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance comparison across different optimization techniques after quantization.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th></th>
<th><inline-formula id="ieqn-267"><mml:math id="mml-ieqn-267"><mml:mi>E</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-268"><mml:math id="mml-ieqn-268"><mml:mi>V</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>F</mml:mi><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-269"><mml:math id="mml-ieqn-269"><mml:mi>C</mml:mi><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>K</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="normal">&#x0025;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-270"><mml:math id="mml-ieqn-270"><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>M</mml:mi><mml:mi>B</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-271"><mml:math id="mml-ieqn-271"><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mtext>&#xA0;</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>MobileNetV3-L (BASE)</italic></td>
<td><inline-formula id="ieqn-272"><mml:math id="mml-ieqn-272"><mml:mn>7.587</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-273"><mml:math id="mml-ieqn-273"><mml:mn>77.512</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-274"><mml:math id="mml-ieqn-274"><mml:mn>67.515</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-275"><mml:math id="mml-ieqn-275"><mml:mn>5.93</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-276"><mml:math id="mml-ieqn-276"><mml:mn>12.7</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Only Attention</italic></td>
<td><inline-formula id="ieqn-277"><mml:math id="mml-ieqn-277"><mml:mn>7.089</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-278"><mml:math id="mml-ieqn-278"><mml:mn>78.590</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-279"><mml:math id="mml-ieqn-279"><mml:mn>67.794</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-280"><mml:math id="mml-ieqn-280"><mml:mn>5.94</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-281"><mml:math id="mml-ieqn-281"><mml:mn>13.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Only Distilled</italic></td>
<td><inline-formula id="ieqn-282"><mml:math id="mml-ieqn-282"><mml:mn>6.417</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-283"><mml:math id="mml-ieqn-283"><mml:mn>81.086</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-284"><mml:math id="mml-ieqn-284"><mml:mn>72.611</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-285"><mml:math id="mml-ieqn-285"><mml:mn>5.93</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-286"><mml:math id="mml-ieqn-286"><mml:mn>12.7</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>Attention &#x0026; Distilled</italic></td>
<td><inline-formula id="ieqn-287"><mml:math id="mml-ieqn-287"><mml:mn>7.173</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-288"><mml:math id="mml-ieqn-288"><mml:mn>79.726</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-289"><mml:math id="mml-ieqn-289"><mml:mn>70.143</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-290"><mml:math id="mml-ieqn-290"><mml:mn>5.94</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-291"><mml:math id="mml-ieqn-291"><mml:mn>13.2</mml:mn></mml:math></inline-formula></td>
</tr>
<tr>
<td><italic>ResNet-50 (SOTA)</italic></td>
<td><inline-formula id="ieqn-292"><mml:math id="mml-ieqn-292"><mml:mn>6.335</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-293"><mml:math id="mml-ieqn-293"><mml:mn>82.271</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-294"><mml:math id="mml-ieqn-294"><mml:mn>75.039</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-295"><mml:math id="mml-ieqn-295"><mml:mn>45.3</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-296"><mml:math id="mml-ieqn-296"><mml:mn>163.5</mml:mn></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this work, post-training quantization is applied to the final models due to its simplicity, compatibility with mobile deployment, and no retraining requirement. Specifically, the weights are converted from float32 to float16, while other tensors remain in full precision due to the model&#x2019;s sensitivity to numerical precision [<xref ref-type="bibr" rid="ref-20">20</xref>]. As float16 naturally preserves the dynamic range of the weights, converting them to half-precision reasonably reduces memory footprint and computational cost while minimizing precision loss.</p>
<p>The baseline model experiences a substantial accuracy drop after quantization, with an EER increasing from <inline-formula id="ieqn-297"><mml:math id="mml-ieqn-297"><mml:mn>1.699</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-298"><mml:math id="mml-ieqn-298"><mml:mn>7.587</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, VR@FAR of <inline-formula id="ieqn-299"><mml:math id="mml-ieqn-299"><mml:mn>77.512</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> (a drop of about <inline-formula id="ieqn-300"><mml:math id="mml-ieqn-300"><mml:mn>20</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), and CMC@K of <inline-formula id="ieqn-301"><mml:math id="mml-ieqn-301"><mml:mn>67.515</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> (a drop of about <inline-formula id="ieqn-302"><mml:math id="mml-ieqn-302"><mml:mn>27</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>), yet it retains a more compact size of <inline-formula id="ieqn-303"><mml:math id="mml-ieqn-303"><mml:mn>5.93</mml:mn></mml:math></inline-formula> MB and lower latency of <inline-formula id="ieqn-304"><mml:math id="mml-ieqn-304"><mml:mn>12.7</mml:mn></mml:math></inline-formula> ms. By comparison, the quantized state-of-the-art ResNet-50 remains more accurate but is significantly larger and slower, which is expected given the typical trade-off between accuracy and computational cost.</p>
<p>Meanwhile, the attention-only and distilled-only variants show slight overall improvements compared to the baseline, showing some resilience to quantization. Importantly, the combined variant achieves a balanced performance among all optimized models, with an EER of <inline-formula id="ieqn-305"><mml:math id="mml-ieqn-305"><mml:mn>7.173</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR of <inline-formula id="ieqn-306"><mml:math id="mml-ieqn-306"><mml:mn>79.726</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K of <inline-formula id="ieqn-307"><mml:math id="mml-ieqn-307"><mml:mn>70.143</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, while maintaining a moderate size of <inline-formula id="ieqn-308"><mml:math id="mml-ieqn-308"><mml:mn>5.94</mml:mn></mml:math></inline-formula> MB and reasonable latency of <inline-formula id="ieqn-309"><mml:math id="mml-ieqn-309"><mml:mn>13.2</mml:mn></mml:math></inline-formula> ms. Overall, this indicates that integrating both enhancements effectively preserves discriminative features and provides a well-balanced trade-off between accuracy and efficiency, demonstrating its practicality for mobile deployment.</p>
<p>In general, the proposed optimizations enhance accuracy and maintain efficiency relative to the baseline; however, all models suffer a noticeable drop in recognition accuracy after quantization, with EER increasing, and VR@FAR and CMC@K decreasing across all variants. To address this, alternative strategies such as quantization-aware training or mixed precision could be considered to reduce the performance drop, though they involve additional training complexity and are left for future work. Consequently, this highlights the need for further development to better preserve accuracy and improve practicality for mobile deployment.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Conclusion</title>
<p>This study presents a lightweight MobileNetV3-L model for mobile iris recognition, optimized through attention mechanism, knowledge distillation, and quantization to achieve a practical balance between accuracy and efficiency. Comprehensive experiments demonstrate that the model achieves a compact size of <inline-formula id="ieqn-310"><mml:math id="mml-ieqn-310"><mml:mn>12</mml:mn></mml:math></inline-formula> MB and a low latency of <inline-formula id="ieqn-311"><mml:math id="mml-ieqn-311"><mml:mn>27</mml:mn></mml:math></inline-formula> ms, facilitating its adaptation for mobile applications through a further reduction to <inline-formula id="ieqn-312"><mml:math id="mml-ieqn-312"><mml:mn>5.94</mml:mn></mml:math></inline-formula> MB and <inline-formula id="ieqn-313"><mml:math id="mml-ieqn-313"><mml:mn>13</mml:mn></mml:math></inline-formula> ms after post-training quantization. Regarding recognition performance, the proposed combined variant, the main model of this study, delivers an EER of <inline-formula id="ieqn-314"><mml:math id="mml-ieqn-314"><mml:mn>1.409</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, a VR@FAR&#x003D;<inline-formula id="ieqn-315"><mml:math id="mml-ieqn-315"><mml:mn>1</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> of <inline-formula id="ieqn-316"><mml:math id="mml-ieqn-316"><mml:mn>98.184</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and a CMC@K&#x003D;<inline-formula id="ieqn-317"><mml:math id="mml-ieqn-317"><mml:mn>1</mml:mn></mml:math></inline-formula> of <inline-formula id="ieqn-318"><mml:math id="mml-ieqn-318"><mml:mn>94.785</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, all meeting the respective targets of <inline-formula id="ieqn-319"><mml:math id="mml-ieqn-319"><mml:mn>5</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-320"><mml:math id="mml-ieqn-320"><mml:mn>95</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-321"><mml:math id="mml-ieqn-321"><mml:mn>90</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula> on the standard CASIA-IrisV4-Thousand dataset. Overall, the results highlight that the proposed strategies enable models to achieve performance comparable to state-of-the-art approaches.</p>
<p>Building on these findings, the practical value of this work lies in providing a mobile-ready iris recognition solution that balances recognition accuracy and computational efficiency, improving the overall system architecture to achieve robust and reliable performance, and demonstrating the framework&#x2019;s practicality through extensive experimental evaluation. By utilizing lightweight architectures and optimization techniques, the proposed framework facilitates deployment across a wide range of practical domains, including smart cities and mobile security. Despite these strengths, current evaluations are limited to standard datasets, and maintaining accuracy under aggressive quantization remains challenging. Consequently, future research built on this work can further expand this work through validation on larger and more diverse datasets, exploration of advanced optimization strategies, extension to multimodal biometric systems, and evaluation across different hardware platforms to enhance robustness, scalability and generalizability.</p>
</sec>
</body>
<back>
<ack>
<p>The authors extend their appreciation to the Posts and Telecommunications Institute of Technology (PTIT, Vietnam) for supporting this research.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>Not applicable.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Trong-Thua Huynh and De-Thu Huynh; methodology, Trong-Thua Huynh and Du-Thang Phu; software: Du-Thang Phu and De-Thu Huynh; validation, Quoc H. Nguyen and Hong-Son Nguyen; formal analysis, Trong-Thua Huynh and Quoc H. Nguyen, resources, Du-Thang Phu; data curation, Trong-Thua Huynh; writing&#x2014;original draft preparation, Du-Thang Phu and De-Thu Huynh; writing&#x2014;review and editing, Trong-Thua Huynh and Du-Thang Phu; visualization, De-Thu Huynh; supervision, Hong-Son Nguyen; project administration, Trong-Thua Huynh. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the Corresponding Author, upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Grother</surname> <given-names>P</given-names></string-name>, <string-name><surname>Matey</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tabassi</surname> <given-names>E</given-names></string-name>, <string-name><surname>Quinn</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chumakov</surname> <given-names>M</given-names></string-name></person-group>. <chapter-title>IREX VI: temporal stability of iris recognition accuracy</chapter-title>. In: <source>NIST Interagency/Internal Report (NISTIR) 7948</source>. <publisher-loc>Gaithersburg, MD, USA</publisher-loc>: <publisher-name>National Institute of Standards and Technology</publisher-name>; <year>2013 Jul 11</year>. doi:<pub-id pub-id-type="doi">10.6028/NIST.IR.7948</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Emergen Research</collab></person-group>. <article-title>Iris recognition market, by component, product, application, end-use, and region, forecast to 2034</article-title>; <year>2025 Jul [cited 2025 Oct 15]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.emergenresearch.com/industry-report/iris-recognition-market">https://www.emergenresearch.com/industry-report/iris-recognition-market</ext-link>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Mordor Intelligence Research &#x0026; Advisory</collab></person-group>. <article-title>Iris recognition market size &#x0026; share analysis: growth trends &#x0026; forecasts (2025&#x2013;2030) [Internet]</article-title>. <year>2025 Jul [cited 2025 Oct 15]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.mordorintelligence.com/industry-reports/iris-recognition-market">https://www.mordorintelligence.com/industry-reports/iris-recognition-market</ext-link>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <comment>arXiv:1409.1556. 2014</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.1409.1556</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Van Der Maaten</surname> <given-names>L</given-names></string-name>, <string-name><surname>Weinberger</surname> <given-names>KQ</given-names></string-name></person-group>. <article-title>Densely connected convolutional networks</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>4700</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Proen&#x00E7;a</surname> <given-names>H</given-names></string-name>, <string-name><surname>Alonso-Fernandez</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Deep learning for iris recognition: a survey</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>9</issue>):<fpage>223:1&#x2013;35</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3651306</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>GhostNet: more features from cheap operations</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <year>2020 Jun 13&#x2013;19</year>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>1580</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00165</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Boutros</surname> <given-names>F</given-names></string-name>, <string-name><surname>Damer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Raja</surname> <given-names>KB</given-names></string-name>, <string-name><surname>Ramachandra</surname> <given-names>R</given-names></string-name>, <string-name><surname>Kirchbuchner</surname> <given-names>F</given-names></string-name>, <string-name><surname>Kuijper</surname> <given-names>A</given-names></string-name></person-group>. <article-title>On benchmarking iris recognition within a head-mounted display for AR/VR applications</article-title>. In: <conf-name>Proceedings of the IEEE Int Joint Conference Biometrics (IJCB); 2020 Sep 28&#x2013;Oct 1</conf-name>; <publisher-loc>Houston, TX, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.1109/IJCB48548.2020.9304919</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sumit</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anavatti</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tahtali</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mirjalili</surname> <given-names>S</given-names></string-name>, <string-name><surname>Turhan</surname> <given-names>U</given-names></string-name></person-group>. <article-title>ResNet-Lite: on improving image classification with a lightweight network</article-title>. <source>Comput Sci</source>. <year>2024</year>;<volume>246</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2024.09.597</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hinton</surname> <given-names>G</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Dean</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Distilling the knowledge in a neural network</article-title>. <comment>arXiv:1503.02531</comment>. <year>2015</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Maybank</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Knowledge distillation: a survey</article-title>. <source>Int J Comput Vis</source>. <year>2021</year>;<volume>129</volume>(<issue>6</issue>):<fpage>1789</fpage>&#x2013;<lpage>819</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-021-01453-z</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Han</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep learning for iris recognition: a review</article-title>. <comment>arXiv:2303.08514</comment>. <year>2023</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Latif</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Sidek</surname> <given-names>KA</given-names></string-name>, <string-name><surname>Bakar</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Hashim</surname> <given-names>AHA</given-names></string-name></person-group>. <article-title>Online multimodal compression using pruning and knowledge distillation for iris recognition</article-title>. <source>J Adv Res Appl Sci Eng Technol</source>. <year>2024</year>;<volume>37</volume>(<issue>2</issue>):<fpage>68</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.37934/araset.37.2.6881</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>Institute of Automation, Chinese Academy of Sciences (CASIA)</collab></person-group>. <source>CASIA Iris Image Database, Version 4.0</source>. <publisher-loc>Beijing, China</publisher-loc>: <publisher-name>Chinese Academy of Sciences</publisher-name>; <year>2010</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Worldcoin</surname> <given-names>AI</given-names></string-name></person-group>. <article-title>IRIS: iris recognition inference system of the worldcoin project [Computer software]</article-title>. <source>GitHub</source>. <year>2023 [cited 2025 Oct 15]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/worldcoin/open-iris">https://github.com/worldcoin/open-iris</ext-link>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zafeiriou</surname> <given-names>S</given-names></string-name></person-group>. <article-title>ArcFace: additive angular margin loss for deep face recognition</article-title>. <comment>arXiv:1801.07698</comment>. <year>2019</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Raj</surname> <given-names>B</given-names></string-name>, <string-name><surname>Song</surname> <given-names>L</given-names></string-name></person-group>. <article-title>SphereFace: deep hypersphere embedding for face recognition</article-title>. <comment>arXiv:1704.08063</comment>. <year>2017</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Large margin cosine loss for deep face recognition</article-title>. <comment>arXiv:1801.09414</comment>. <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jacob</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kligys</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Howard</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Quantization and training of neural networks for efficient integer-arithmetic-only inference</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>2704</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00286</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>