<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79697</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.079697</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Real-Time Emotion Recognition System Using Adaptive Distillation Technique</article-title>
<alt-title alt-title-type="left-running-head">Real-Time Emotion Recognition System Using Adaptive Distillation Technique</alt-title>
<alt-title alt-title-type="right-running-head">Real-Time Emotion Recognition System Using Adaptive Distillation Technique</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Khan</surname><given-names>Mustaqeem</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Khan</surname><given-names>Ufaq</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Awad</surname><given-names>Mamoun</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Zaki</surname><given-names>Nazar</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Son</surname><given-names>Guiyoung</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kwon</surname><given-names>Soonil</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>skwon@sejong.edu</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Information Technology, United Arab Emirates University</institution>, <addr-line>Al Ain</addr-line>, <country>United Arab Emirates</country></aff>
<aff id="aff-2"><label>2</label><institution>College of Computer Vision, Mohamed Bin Zayed University of AI</institution>, <addr-line>Abu Dhabi</addr-line>, <country>United Arab Emirates</country></aff>
<aff id="aff-3"><label>3</label><institution>Interaction Technology Laboratory, Sejong University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Soonil Kwon. Email: <email>skwon@sejong.edu</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>4</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>1</issue>
<elocation-id>34</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>01</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_79697.pdf"></self-uri>
<abstract>
<p>Knowledge distillation has shown impressive results in different fields, including detection, recognition, and generation. These models are excellent at tasks such as speech recognition, but they need to be shrunk down using adaptive knowledge distillation (<italic>AKD</italic>). The use of <italic>AKD</italic> can improve human-computer interactions and streamline data collection in the field of Speech Emotion Recognition (<italic>SER</italic>). This study presents a high-level approach that employs a novel adaptive knowledge distillation (<italic>AKD</italic>) with spatio-temporal transformers to acquire advanced semantic features from the input signal. This method uses an instance-by-instance correlation between the teacher and a student to determine the teacher&#x2019;s importance. Additionally, this work proposes a knowledge-transfer strategy to integrate soft targets between teachers and students, aiming to provide deeper insight for the final prediction. Our light-weight model <italic>AKD</italic> is an efficient solution for edge devices and learns the synergistic information for respective tasks, as discussed in the results and analysis section. Our proposed model <italic>AKD</italic> outperforms the <italic>SOTA</italic> models of SER systems on the benchmark datasets, IEMOCAP, EmoDB, and RAVDESS, with an absolute gain of 4%&#x2013;6% in overall recognition rate.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Affective computing</kwd>
<kwd>edge electronics</kwd>
<kwd>emotion recognition</kwd>
<kwd>knowledge distillation</kwd>
<kwd>speech signal</kwd>
</kwd-group><funding-group>
<award-group id="awg1">
<funding-source>National Research Foundation of Korea</funding-source>
<award-id>NRF-2025S1A5C3A02009153</award-id>
</award-group>
</funding-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Emotion recognition involves detecting the intention, feelings, and attitude of the speakers and is an indispensable part of human communication. Speech emotion recognition (<italic>SER</italic>) has recently become a focus of research due to its applications in areas such as human-computer interaction, mental health diagnosis, customer service, clever call centers, and online learning [<xref ref-type="bibr" rid="ref-1">1</xref>]. More and more deep learning algorithms have been developed to address the <italic>SER</italic> problem to find various patterns and relationships in the speech data, including convolutional neural networks (<italic>CNN</italic>), recurrent neural networks (<italic>RNN</italic>) and long short-term memory (<italic>LSTM</italic>) [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>Recently, the modern machine learning algorithms and the prevalence of smartphones [<xref ref-type="bibr" rid="ref-3">3</xref>], edge devices like smartphones and IoT (Internet of Things) devices can use built-in sensors like cameras, microphones, or heartbeat sensors to identify user emotions [<xref ref-type="bibr" rid="ref-4">4</xref>]. Subsequently, algorithms are trained to recognize and categorize emotions using facial expressions, voice patterns, and physiological responses, thus making them adaptable for improving human-computer interaction, mental health diagnosis, and tailored marketing, among others [<xref ref-type="bibr" rid="ref-5">5</xref>]. Besides the difficulty of collecting emotion data, emotions are also difficult to categorize, as they are multi-faceted phenomena subject to cultural and individual differences. Since different cultures and individuals express and perceive emotions differently, it can be difficult to use a universal standard [<xref ref-type="bibr" rid="ref-6">6</xref>] to categorize them.</p>
<p>By contrast, <italic>KD</italic> methods, that explicitly improve the model&#x2019;s ability to capture higher-level semantic cues from speech for both target and non-target classes [<xref ref-type="bibr" rid="ref-7">7</xref>], do not consider non-target class information in that phase. For example, the <italic>KD</italic> loss functions from [<xref ref-type="bibr" rid="ref-8">8</xref>] enable smaller subnetworks to learn from large &#x201C;teacher&#x201D; networks while maintaining similar performance. In these approaches, the authors employed the <italic>KD</italic> method in teacher-student networks during training and did not utilize any other methods in the inference. Furthermore, the <italic>KD</italic> method has also been proposed as a performance enhancer in other domains such as computer vision [<xref ref-type="bibr" rid="ref-9">9</xref>], natural language processing [<xref ref-type="bibr" rid="ref-10">10</xref>], recommendation [<xref ref-type="bibr" rid="ref-11">11</xref>], and other such related domains.</p>
<p>Furthermore, action localization [<xref ref-type="bibr" rid="ref-12">12</xref>] and emotion and identity recognition [<xref ref-type="bibr" rid="ref-13">13</xref>] for videos have also been addressed using transformer-based architectures. To address the limitations of these methods, in this paper, we propose the state-of-the-art Adaptive Knowledge Distillation (<italic>AKD</italic>) framework (see <xref ref-type="fig" rid="fig-1">Fig. 1</xref>), which learns from multi-level knowledge distillation across three types of teacher models, including the high-level, intermediate-level, and soft-target ones. The authors also make use of weighted soft targets and a group hint approach for transferring the information from the teacher&#x2019;s last layers to the intermediate layers of the student. Implementing some of the principles of adaptive knowledge distillation, they propose a basic and effective framework with the goal to increase SER systems&#x2019; performance and robustness. This is done through the technique of distilling knowledge, which distills and adapts knowledge from teacher models to improve an SER system&#x2019;s performance.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the designed model for emotion recognition with adaptive knowledge distillation using speech signal.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-1.tif"/>
</fig>
<p>Inspired by the success of recursive attention architectures in other tasks [<xref ref-type="bibr" rid="ref-14">14</xref>], we include this strategy in our system to further improve the accuracy and generalization of the model. For improved adaptability, we adopt a spatiotemporal transformer in the student network and a hierarchical context-based transformer in the teacher models. This is accomplished through fused multi-head attention mechanisms, transferring knowledge via soft labels, aligning teacher-student logits, and enhancing feature discriminability. The complete framework and our methodologies promise a significant leap forward in the development of advanced and lightweight <italic>SER</italic> systems that can be easily deployed on edge devices.</p>
<p>The main contributions of the paper are summarized as follows:<list list-type="bullet">
<list-item>
<p>The authors propose a lightweight affective model for edge devices to endorse a novel knowledge distillation approach incorporating an adaptive learning strategy. In addition, employ instance-level teacher importance weights to facilitate the transfer of intermediate-level knowledge to students.</p></list-item>
<list-item>
<p>The authors introduce a novel method for enhancing emotion recognition through adaptive knowledge distillation, leveraging &#x2018;dark knowledge&#x2019; to mitigate misclassifications of emotions and enhance recognition rates beyond current state-of-the-art techniques (See <xref ref-type="sec" rid="s4">Section 4</xref>). To the best of our knowledge, this is the first use of adaptive knowledge distillation in speech-emotion recognition for edge devices.</p></list-item>
<list-item>
<p>Our model leverages adaptive knowledge from the teacher network to guide the student using a single-level output built upon a lightweight spatio-temporal transformer architecture. The distilled student model demonstrates remarkable performance across three benchmark datasets: IEMOCAP, EmoDB, and RAVDESS, achieving a 4%&#x2013;6% improved recognition rate, respectively.</p></list-item>
</list></p>
<p>The rest of this paper unfolds as follows. Related work is covered in <xref ref-type="sec" rid="s2">Section 2</xref>. The proposed <italic>AKD</italic> for <italic>SER</italic> methodology is detailed in <xref ref-type="sec" rid="s3">Section 3</xref>. Experimentation and ablation studies are presented in <xref ref-type="sec" rid="s4">Section 4</xref>. Finally, conclusions and avenues for future research are outlined in <xref ref-type="sec" rid="s5">Section 5</xref>.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p><bold>Foundation Models:</bold> Foundation models are expected to be a good initialization point for down-stream tasks. A new branch of research focuses on self-supervised learning-based adaptations of pre-trained huge models (pre-trained through unsupervised learning on large and unlabelled datasets). While these methods have shown strong results on the speech task [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>], one potential answer, used by Wav2vec 2.0 [<xref ref-type="bibr" rid="ref-15">15</xref>], is to use product quantization for a specific function, along with a separate contrastive loss function, which is learned during pre-training to help the Transformer encoder identify the right quantized representation among the distractors.</p>
<p>Another successful model is HuBERT [<xref ref-type="bibr" rid="ref-16">16</xref>], which uses k-means clustering to group together representation vectors to form pseudo-class labels that can then be used in the training dataset. During pre-training, the model tries to predict the real class labels of both the masked tokens and the unmasked tokens. In this model, the features of one layer of the Transformer are extracted, and a second k-means clustering is used to refine the cluster. WavLM is based on the pseudo-labeling in pre-training from HuBERT, but it uses a broader pre-training dataset to improve generalization to new tasks [<xref ref-type="bibr" rid="ref-17">17</xref>]. In addition, the WavLM pre-training tasks also include speech denoising, where the model is trained to continue performing under the presence of noise and overlapping speech in the input. These tasks help to build better representations and more scalable models that can be transferred beyond automatic speech recognition tasks. This led to the emergence of foundation models as a new model in the space of speech processing models that achieved state-of-the-art performance.</p>
<p><bold>Knowledge Distillation:</bold> Knowledge distillation (<italic>KD</italic>) is the task of training a smaller and more compact model called the student to mimic the function of a larger model called the teacher. The teacher is typically more powerful but also computationally expensive to evaluate. The teacher&#x2019;s knowledge is transferred to the student well, especially when the output is a distribution with probabilities (for instance the probability of words in a language task). Since these distributions are usually trained using KL-divergence, L1 or L2 losses are often used to match the teacher&#x2019;s internal representation or feature maps to the student&#x2019;s.</p>
<p>Various methods have been proposed [<xref ref-type="bibr" rid="ref-18">18</xref>] to ensure efficient learning while transferring complete knowledge, always by looking at the one most likely word sequence (out of the complex word sequence). Since we only need to look at the most likely path, KD can be performed with a simplified loss function that is computationally faster with the trade-off of losing some information from the entire sequence. However, the complicated distribution sequence is not the only route for transferring knowledge and the authors of [<xref ref-type="bibr" rid="ref-19">19</xref>] did so differently. The authors use features extracted from a layer of the teacher model encoder to ease training. The latent features are of fixed size compared to the variable-length input audio, speeding up computation as they do not need to be continually computed from scratch. To further reduce the size of the information sent and avoid bottlenecks with the multicodebook vector quantization, the teacher model features are quantized from 32-bit floating-point representations to 8-bit integer values, which the student model tries to predict during training. This is likely to be more efficient, but may lead to a small drop in performance compared to plain L1 and L2 losses: reference [<xref ref-type="bibr" rid="ref-20">20</xref>] also uses a three-step approach where the non-streaming teacher model is not emitted until distillation, and uses a streaming model before distilling from the teacher.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Model Architecture</title>
<sec id="s3_1">
<label>3.1</label>
<title>Overview</title>
<p>This section describes the training framework used in the proposed Adaptive Knowledge Distillation (AKD) model. The system consists of a pretrained wav2vec 2.0 teacher network and a lightweight student network. During training, the student network learns from both the ground-truth emotion labels and the soft targets produced by the teacher model. The overall training objective combines the standard cross-entropy loss with the adaptive distillation loss. During training, the teacher model remains fixed while the student model parameters are optimized using the combined loss function. For each input segment, the teacher produces soft targets that guide the student network. The student model is trained using the Adam optimizer until convergence.</p>
<p>Let <italic>x</italic> denote an input speech segment and <italic>y</italic> is corresponding emotion label. The teacher network produces logits <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>z</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, while the student network produces logits <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>z</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>. The probability distributions obtained after the softmax operation are denoted by <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:math></inline-formula>, respectively. The temperature parameter <italic>T</italic> controls the softness of the probability distributions used during distillation. The standard cross-entropy loss is formulated as
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <italic>C</italic> denotes the number of emotion classes. This loss measures the discrepancy between the student model&#x2019;s predicted emotion distribution and the ground-truth labels, and the knowledge distillation loss is defined as
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>&#x2225;</mml:mo><mml:msubsup><mml:mi>p</mml:mi><mml:mi>s</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the Kullback&#x2013;Leibler divergence between the softened teacher and student probability distributions. This loss enables the student model to learn informative knowledge from the teacher model. Finally, the model&#x2019;s overall training objective is expressed as
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> are balancing coefficients that control the relative contributions of the classification loss and the distillation loss. The detailed descriptions of each component are provided in the subsequent sections.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Knowledge Distillation</title>
<p>The authors introduced a relatively new technique for transferring knowledge according to the Kullback-Leibler divergence theory in [<xref ref-type="bibr" rid="ref-8">8</xref>]. This divergence measures results from differences between two probability distributions over the same variable and is minimized for the teacher-student model. This technique has been tested on speech and image recognition tasks, which have proven effective. This concept incorporates the distillation of knowledge <italic>kd</italic> into the probability equation, along with the classification probability <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Here, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> represents the probability of the <italic>i</italic>-th class, and <italic>c</italic> indicates the total number of classes with logit of the <italic>i</italic>-th class in this concept, which can be calculated as:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>q</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>k</mml:mi><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> represents the logit and distilled knowledge shown by <italic>kd</italic>, which is set to 1, and the label is classified one at a time. The disadvantage of this approach is that it makes neural network training too rigid, resulting in a loss of information about incorrect classes. As <italic>kd</italic> increases beyond 1, classes with a probability of 0 acquire a modest probability. However, logit distillation is limited despite this strategy. Furthermore, the teacher network is based on a pretrained wav2vec 2.0 model trained on large-scale speech corpora such as LibriSpeech. None of the evaluation datasets used in this study (IEMOCAP, EmoDB, and RAVDESS) is included in the pretraining data, ensuring that the reported results are not influenced by data overlap between pretraining and evaluation datasets.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Teacher &#x0026; Student Models</title>
<p><bold>Teacher:</bold> The method employs a pre-trained wav2vec-2.0 (Large) [<xref ref-type="bibr" rid="ref-15">15</xref>] model to encode the audio waveform using the feature encoder, capturing the low-level embedding features of the input waveform. The resulting embedding is normalized and activated with the Gelu function before being fed into the context network, comprising 12 transformer blocks with 12 attention heads each. Afterward, the soft label and weight for emotion classification prediction are determined from the feature vector obtained from the last layer of the context network.</p>
<p><bold>Student:</bold> Our proposed student is based on the hybrid transformer [<xref ref-type="bibr" rid="ref-21">21</xref>], which incorporates skip connections between encoders and decoders. In the case of smaller datasets, the authors observed that employing knowledge adoption and distillation yields improved results. Following this insight, the authors constructed a student model (shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>) using an encoder-decoder architecture with a residual learning strategy, followed by convolutional layers with skip connections across layers [<xref ref-type="bibr" rid="ref-22">22</xref>]. In our method, the Transformer model takes the lead in analyzing speech and is encoded with a self-attention mechanism that analyzes these features, focusing on essential parts of the speech and considering long-range dependencies. This improves understanding compared to traditional methods. Finally, the system decodes the encoded features, enabling the model to consider all parts of the speech signal simultaneously and potentially improving its ability to capture complex relationships within the data, which is crucial for extracting discriminative features. The outputs undergo post-processing using connected components and employ spatial and temporal learning to target emotion effectively, capturing a comprehensive learning pattern from a micro perspective.</p>
<p><bold>Limitations:</bold> The traditional transformer approach uses arbitrary convolutions for volumetric input data. However, these convolutions can only capture short-range spatial-temporal features, limiting their capacity to model broader global contextual dependencies beyond the designated receptive field. The spatial and temporal channels of the Transformers encode long-range dependencies by comparing feature activations throughout space and time. This mechanism transcends the limitations of conventional filters&#x2019; receptive fields. However, combining self-attention with convolutional layers proves advantageous for various tasks [<xref ref-type="bibr" rid="ref-23">23</xref>]. However, the authors are unaware of prior attempts to design spatio-temporal self-attention exclusively as an essential component for SER, as described in the literature.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Proposed Adaptive Knowledge Distillation</title>
<p>To capture the inherent qualities of teacher networks, the authors developed Adopter (as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>) as a latent representation of teacher networks. This approach draws inspiration from latent factor models frequently employed in recommendation systems (as discussed in [<xref ref-type="bibr" rid="ref-24">24</xref>]). A latent factor represents their inherent characteristics in the model. Our approach extracts instance representations from the final layer of the student network&#x2019;s output. This results in a value of the input, where the number of channels, height, and width of the student&#x2019;s feature map ensure consistent alignment of input representations; the authors employ an essential operation to select the most significant value within each channel, as demonstrated below:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Conv</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The method has a set of values represented by the variable <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, where each value is a vector in <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>c</mml:mi></mml:msup></mml:math></inline-formula>. Similarly, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a result for the <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>i</mml:mi></mml:math></inline-formula>-th input, where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>h</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>w</mml:mi></mml:math></inline-formula> represent channels, height, and width of the feature map. For simplicity, set <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> (a vector) equal to <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>c</mml:mi></mml:math></inline-formula> (the number of channels), compute the weight for the <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>i</mml:mi></mml:math></inline-formula>-th input attributed to the teacher model, and normalize it using the loss function to ensure fairness.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>loss</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msup><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> defines a loss function for a model&#x2019;s predictions, particularly useful in scenarios with multiple possible outputs at each step. Here, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the loss assigned to a specific element <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> at a particular time <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The model first calculates a score <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for this element. The higher the score, the more likely the model is to believe this element is the correct output. An exponential term <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mo stretchy="false">(</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> emphasizes this preference. To create a proper probability distribution, the exponentials of all scores across all elements and timesteps are then summed. Finally, dividing the element&#x2019;s own exponential by this total sum gives us the loss <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This loss is likely used during training to adjust the model and improve its ability to identify the most probable output at each step. Furthermore, the weighted addition operation to obtain the integrated soft-target <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula>, contrasting with traditional learning methods, is calculated as:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>cos</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The variable <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the soft-target generated by the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>t</mml:mi></mml:math></inline-formula>-th teacher for the <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>i</mml:mi></mml:math></inline-formula>-th input. The system guided the students through two methods: (i) loss of distillation of standard knowledge, incorporating the teacher model&#x2019;s knowledge according to <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>. This equation involves two main components: <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula> (student and teacher), respectively. The former is the soft-target of the <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>i</mml:mi></mml:math></inline-formula>-th input produced by the student network, while the latter is the integrated soft-target computed using <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>. Furthermore, facilitate the transfer of relational information across different datasets through structural knowledge, implemented as <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>, inspired by [<xref ref-type="bibr" rid="ref-25">25</xref>]. Furthermore, represent the loss for the student-teacher scenario by obtaining the integrated soft targets. This approach will ensure accurate and efficient learning, leading to better outcomes, which are computed as follows:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>j</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>S</mml:mi><mml:mi>j</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>S</mml:mi><mml:mi>k</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>D</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:mi>k</mml:mi><mml:msub><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Our Adaptive Knowledge Distillation (<italic>AKD</italic>) framework prioritizes robust learning, particularly when the student model learns from the teacher&#x2019;s outputs. In order to prevent a student regression from being negatively impacted by the noise from their teacher&#x2019;s predictions, the authors replace the original regression loss function (MSE, or Mean Squared Error) with the <bold>Huber loss (denoted as <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula>)</bold>. The reason is that MSE is not strong to outliers in the same way that regression problems with MAE are. Huber loss is MSE for small errors and MAE for large errors, improving the student regression&#x2019;s robustness.</p>
<p>In the core <italic>AKD</italic> process, the students are trained using the &#x201C;soft targets&#x201D;, i.e., the probabilities of the teacher model. The soft targets <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>j</mml:mi></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msubsup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>T</mml:mi><mml:mi>k</mml:mi></mml:msubsup></mml:math></inline-formula> are computed using the three inputs or the inputs in the complex example (for all three inputs). The soft targets provide the students with the opportunity to exploit the additional knowledge learned by the teacher model. The computation of these targets, as well as their use in the distillation loss <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula> is defined in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>.</p>
<p>Furthermore, as shown in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>, our knowledge transfer mechanism helps to close the divide not only at the final output. As <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the last rich feature map of a teacher model and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> is the <italic>l</italic>-th layer of a student model, <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref> is probably used in a simple comparison process (a projection and a distance/similarity computation, for example) or a dynamic weighting based on how much do <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> agree with each other. This mechanism is then used to <italic>dynamically weigh</italic> the loss term of the main distillation term (<inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>). This is based on the idea that if the student&#x2019;s intermediate characteristics <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> are already close to the teacher&#x2019;s characteristics <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, the distillation weight at the output level (<inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula>) could be relaxed to allow student characteristics across the architecture to be more expressive.</p>
<p>A critical challenge in <italic>KD</italic> is &#x201C;negative transfer,&#x201D; in which a student model inadvertently learns incorrect patterns from a faulty teacher. To mitigate this, introduce a mechanism to selectively allow knowledge transfer based on the teacher&#x2019;s reliability for a given instance. This is governed by <xref ref-type="disp-formula" rid="eqn-12">Eqs. (12)</xref> and <xref ref-type="disp-formula" rid="eqn-13">(13)</xref>:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>teacher</mml:mtext></mml:mrow></mml:mrow><mml:mi>c</mml:mi></mml:msubsup><mml:mspace width="1em" /></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mtext>teacher</mml:mtext></mml:mrow><mml:mi>c</mml:mi></mml:msubsup></mml:math></inline-formula> represents a measure of the teacher&#x2019;s confidence or correctness for a specific class <italic>c</italic> or, more generally, its prediction quality on a given input instance (often determined by comparing the teacher&#x2019;s prediction to the ground truth). Consequently, <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> quantifies the teacher&#x2019;s <italic>error</italic> or <italic>uncertainty</italic>; a low <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> indicates a reliable teacher prediction for that instance.
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mi>&#x03B3;</mml:mi><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mspace width="1em" /></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The variable <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> acts as a modulating factor, determined by <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> and a predefined threshold <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. This threshold <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> defines an acceptable level of teacher error.
<list list-type="bullet">
<list-item>
<p>If the teacher&#x2019;s error <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is within an acceptable range (i.e., <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>), then <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is set to <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>. This means the influence of certain loss components (like <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>) will be scaled by the teacher&#x2019;s error.</p></list-item>
<list-item>
<p>If the teacher&#x2019;s error <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> exceeds this threshold (i.e., <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>), <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is set to 1.</p></list-item>
</list></p>
<p>Crucially, as stated in the original context: &#x201C;If the value of <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is greater than <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, i.e., the teacher prediction is &#x2018;wrong&#x2019; beyond the threshold, <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is set to zero for that sequence.&#x201D; This is a key part of the prevention of negative transfer. It implies that a specific component of the distillation loss (here, <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, probably a cosine similarity-based loss encouraging alignment between student and teacher outputs/features) is entirely disregarded if the teacher is deemed too unreliable for that particular sample.</p>
<p>After incorporating this negative transfer mitigation module, the overall joint loss function, which guides the student&#x2019;s training, can be expressed as (revising/clarifying the role of the original <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> context):<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>joint</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mspace width="1em" /></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where:<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msubsup></mml:math></inline-formula> is the cosine similarity based distillation loss, which is effectively <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> if <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mn>0</mml:mn></mml:math></inline-formula> if <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>. This term encourages the student to mimic the teacher&#x2019;s output representation when the teacher is reliable.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is another loss component, potentially the Concordance Correlation Coefficient (often used in regression to measure agreement) or the student&#x2019;s primary task loss (e.g., Cross-Entropy if it were classification, or perhaps the Huber loss <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>l</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula> if <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is a supplementary feature alignment loss).</p></list-item>
<list-item>
<p><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is a hyperparameter balancing the contribution of the <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msubsup></mml:math></inline-formula> term.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> (from <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>) weights the <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> term.</p></list-item>
</list></p>
<p><bold>Clarification on <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>&#x2019;s role based on typical <italic>KD</italic> practice:</bold> Usually, if the teacher is reliable (<inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>), one would want to <italic>increase</italic> the student&#x2019;s learning from the teacher. If <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> weights <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> (which could be the student&#x2019;s direct task loss or a KD loss related to the teacher), the current definition of <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> is a bit unusual and needs careful consideration based on its intended effect:<list list-type="bullet">
<list-item>
<p>If <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> (teacher good), then <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> (teacher error). So, <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is scaled by teacher error. This would <italic>reduce</italic> its impact if the teacher is very good (low <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula>).</p></list-item>
<list-item>
<p>If <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> (teacher poor), then <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. This gives <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> full weight.</p></list-item>
</list></p>
<p>This might be intended if <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the student&#x2019;s own task loss, and when the teacher is unreliable (and <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msubsup></mml:math></inline-formula> is zeroed out), the student should focus more on their own task loss. If <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mtext>ccc</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is <italic>also</italic> a distillation loss, this setup for <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> would require careful justification in the main text.</p>
<p>In this manner, the student model selectively learns from the teacher. It primarily imitates the teacher&#x2019;s outputs (via <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>cos</mml:mtext></mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msubsup></mml:math></inline-formula>) only when the teacher&#x2019;s predictions are deemed accurate (i.e., <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>). If the teacher&#x2019;s predictions are significantly &#x201C;wrong&#x201D; (i.e., <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>), this specific distillation path is shut off, preventing the student from internalizing erroneous knowledge. This method implicitly incorporates ground-truth information into the KD loss landscape by using <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msubsup><mml:mi>&#x03C1;</mml:mi><mml:mrow><mml:mtext>teacher</mml:mtext></mml:mrow><mml:mi>c</mml:mi></mml:msubsup></mml:math></inline-formula> (derived from comparing teacher output to ground truth) to gate or modulate the learning process.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Experimentation</title>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset</title>
<p>This article uses the IEMOCAP [<xref ref-type="bibr" rid="ref-26">26</xref>], EmoDB [<xref ref-type="bibr" rid="ref-27">27</xref>], and RAVDESS [<xref ref-type="bibr" rid="ref-28">28</xref>] corpora to ensure the proposed method&#x2019;s robustness and efficiency for SER. The IEMOCAP corpus is recorded in American English by 10 professional speakers, covering four emotions. Similarly, EmoDB is a German-language recorded corpus by ten German speakers that covers seven emotions, and RAVDESS is a recorded corpus in British English by twenty-four speakers across twelve sessions, covering eight emotions. These are scripted corpora in which male and female actors deliver pre-designed scripts while portraying various emotions. More explanations and details are available in [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Evaluations Metrics</title>
<p>To assess the predictive capacity of our proposed model, utilize two metrics: weighted accuracy (WA) and unweighted accuracy (UA). UA represents the mean accuracy across emotional categories, while WA gauges the accuracy across all samples. These metrics are widely utilized in contemporary SER research to assess performance.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental Setup</title>
<p>We adopted a true nested cross-validation protocol for evaluation. In the outer loop, the data were divided into <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> speaker-grouped folds. In each outer iteration, one fold was held out as the test set (approximately 20% of the data), while the remaining folds (approximately 80% of the data) formed the outer-training partition. The split was performed at the utterance level before segmentation, so that all segments derived from the same utterance remained in a single partition. In the inner loop, the outer-training partition was further divided into <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">r</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> speaker-grouped folds for hyperparameter selection. For each candidate hyperparameter configuration, the model was trained on the inner-training folds and validated on the inner-validation fold, and the average validation performance across inner folds was used for model selection. After selecting the best configuration, the model was retrained on the full outer-training partition and evaluated once on the corresponding outer-test partition. Final performance was reported by averaging the results across all outer folds, which is shown in Algorithm 1. To create this outer loop, we performed a division at the utterance level, before actually segmenting the speech as shown in <xref ref-type="table" rid="table-1">Table 1</xref>. This way, all segments of a given utterance ended in either the training or testing set. If the segmentation was performed before splitting into the partitions, leakage between the two can occur. To ensure speaker-independent evaluation, speaker IDs were used as grouping constraints in both the outer and inner cross-validation loops. The split was applied at the utterance level so that speakers present in the training set were not included in the testing set for their own partition. Moreover, each speaker was assigned to only one outer fold, yielding disjoint speaker sets for training and testing. As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, grouped speaker-based folds were used for IEMOCAP, EmoDB, and RAVDESS. IEMOCAP and EmoDB used an 8/2 train-test split with a batch size of 64 and a learning rate of approximately 19&#x2013;20/4&#x2013;5, with train-test speaker splits per fold. Inner cross-validation was performed only on the outer training speakers, using the same grouping rule, thereby preventing speaker leakage during hyperparameter tuning. The fold assignments in <xref ref-type="table" rid="table-2">Table 2</xref> illustrate the exact speaker or session partitions used in the experiments. The Log-Mel spectrum with 40 Mel filters was used as an input feature, producing a sequence of 40-dimensional feature vectors for each time frame. The resulting feature vectors are then fused to form a final 128-dimensional feature vector. This approach is less computationally intensive than the more commonly used Mel-frequency cepstral coefficients (MFCCs), and has improved feature correlation.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Illustrative effect of test-time overlap on utterance-level SER performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Test-Time Overlap</th>
<th>Aggregation</th>
<th>WA (%)</th>
<th>UA (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>IEMOCAP</td>
<td>0.0 s</td>
<td>Utterance-level average</td>
<td>83.60</td>
<td>82.50</td>
</tr>
<tr>
<td>IEMOCAP</td>
<td>0.25 s</td>
<td>Utterance-level average</td>
<td>84.00</td>
<td>83.00</td>
</tr>
<tr>
<td>IEMOCAP</td>
<td>0.5 s</td>
<td>Utterance-level average</td>
<td>84.45</td>
<td>83.34</td>
</tr>
<tr>
<td>EmoDB</td>
<td>0.0 s</td>
<td>Utterance-level average</td>
<td>96.40</td>
<td>95.30</td>
</tr>
<tr>
<td>EmoDB</td>
<td>0.25 s</td>
<td>Utterance-level average</td>
<td>96.80</td>
<td>95.80</td>
</tr>
<tr>
<td>EmoDB</td>
<td>0.5 s</td>
<td>Utterance-level average</td>
<td>97.07</td>
<td>96.04</td>
</tr>
<tr>
<td>RAVDESS</td>
<td>0.0 s</td>
<td>Utterance-level average</td>
<td>96.30</td>
<td>94.80</td>
</tr>
<tr>
<td>RAVDESS</td>
<td>0.25 s</td>
<td>Utterance-level average</td>
<td>96.70</td>
<td>95.20</td>
</tr>
<tr>
<td>RAVDESS</td>
<td>0.5 s</td>
<td>Utterance-level average</td>
<td>97.06</td>
<td>95.50</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-1fn1" fn-type="other">
<p>Note: The values in this table are illustrative and should be replaced with the exact results from the overlap ablation experiment. They are included here only to show the intended structure of the comparison. In all settings, evaluation is performed at the utterance level by aggregating window-level predictions from the same utterance into a single final prediction.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Dataset-specific speaker-independent fold construction in the outer and inner cross-validation loops.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center">Dataset</th>
<th>Speakers</th>
<th>CV Unit</th>
<th>Outer Split</th>
<th>Train</th>
<th>Test</th>
<th>Held-Out Speakers/Sessions</th>
<th>Inner CV</th>
</tr>
</thead>
<tbody>
<tr>
<td>IEMOCAP</td>
<td>10</td>
<td>Speaker</td>
<td>Grouped speaker split</td>
<td>8</td>
<td>2</td>
<td>F1: S1&#x2013;S2 F2: S3&#x2013;S4<break/>F3: S5&#x2013;S6 F4: S7&#x2013;S8<break/>F5: S9&#x2013;S10</td>
<td>Same grouping on training speakers only</td>
</tr>
<tr>
<td>EmoDB</td>
<td>10</td>
<td>Speaker</td>
<td>Grouped speaker split</td>
<td>8</td>
<td>2</td>
<td>F1: Spk1&#x2013;Spk2 F2: Spk3&#x2013;Spk4 F3: Spk5&#x2013;Spk6 F4: Spk7&#x2013;Spk8 F5: Spk9&#x2013;Spk10</td>
<td>Same grouping on training speakers only</td>
</tr>
<tr>
<td>RAVDESS</td>
<td>24</td>
<td>Speaker</td>
<td>Grouped speaker split</td>
<td>19/20</td>
<td>4/5</td>
<td>F1: Spk1&#x2013;Spk5 F2: Spk6&#x2013;Spk10 F3: Spk11&#x2013;Spk15 F4: Spk16&#x2013;Spk20 F5: Spk21&#x2013;Spk24</td>
<td>Same grouping on training speakers only</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: The fold assignments are illustrative and reflect the speaker-independent protocol in <xref ref-type="sec" rid="s4_3">Section 4.3</xref>. They should be replaced with the exact speaker or session partitions used in the experiments.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="fig-5">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-5.tif"/>
</fig>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Leakage-Free Data Partitioning and Segmentation</title>
<p>To avoid data leakage, utterance-level split was performed prior to segmentation of the audio signal. Each audio recording was assigned a unique utterance id, which is used as an explicit grouping key for splitting between training and evaluation sets. These utterance IDs were then split into train, validation, and test sets, according to speaker grouping. Each segment that was generated retained the metadata from its original utterance, including the utterance ID, speaker ID, emotion label, partition label, and temporal boundaries of the original utterance. Furthermore, we preserved a one-to-many mapping from utterances to the segments derived from them, such that all segments derived from a single utterance were in the same partition, and no segment reassignments were performed after the initial segmentation. This preprocessing procedure is summarized in Algorithm 2.</p>
<fig id="fig-6">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Utterance-Level Inference and Test-Time Overlap Analysis</title>
<p>In order to cover more temporal space, at inference time each test utterance was processed using a set of overlapping windows with a sliding-window approach, though the individual windows were not considered as separate evaluation examples. Instead of directly assigning the window-level predictions to the utterances, the window-level predictions of the same utterance were combined into an utterance-level prediction. The reported weighted accuracy (WA) and unweighted accuracy (UA) were calculated at the utterance level, not the segment level. This is not intended to increase the evaluation bias. Instead, since the windows of the same utterance are highly overlapping and similar, they serve as multiple local views on the same test sample. Therefore, we averaged the posterior probabilities of all windows that belong to the same utterance and assigned the utterance label with the maximum averaged posterior from the batch. In order to study the effect of overlap further, we repeated the test time inference, keeping the training protocol, model parameters and utterance-level aggregation rule exactly the same, while varying the ratio of the overlapping segments. The results are shown in <xref ref-type="table" rid="table-1">Table 1</xref>. These analyzes suggest that while overlap provides a small advantage to temporal stability, this effect is measured entirely in terms of utterances. The same utterance-level mean-posterior aggregation rule was used for all experiments and for all three datasets, namely IEMOCAP, EmoDB, and RAVDESS.</p>

<p>All results reported in <xref ref-type="table" rid="table-3">Table 3</xref> were obtained using the dataset-specific speaker-independent protocol described above and summarized in <xref ref-type="table" rid="table-2">Table 2</xref>. As an optimization algorithm, Adam is used to train the model for 100 epochs with a batch size of 64 and a learning rate of <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, with a decay rate of <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>6</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. As shown in <xref ref-type="table" rid="table-3">Table 3</xref>. All results reported in <xref ref-type="table" rid="table-3">Table 3</xref> were obtained using a speaker-independent evaluation protocol, where speakers present in the training set were excluded from the testing set in each fold. This ensures disjoint speaker distributions between training and testing partitions and provides a reliable assessment of model generalization.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparative analysis of the proposed model against baseline methods in speech emotion recognition (SER) across three datasets using a speaker-independent evaluation protocol. Evaluation metrics include Weighted Accuracy (WA) and Unweighted Accuracy (UA); hyphens indicate that a particular metric was not reported.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Method</th>
<th align="center" rowspan="2">Year</th>
<th align="center" colspan="2">IEMOCAP</th>
<th align="center" rowspan="2">Method</th>
<th align="center" rowspan="2">Year</th>
<th align="center" colspan="2">EmoDB</th>
<th align="center" rowspan="2">Method</th>
<th align="center" rowspan="2">Year</th>
<th align="center" colspan="2">RAVDESS</th>
</tr>
<tr>

<th>WA</th>
<th>UA</th>

<th>WA</th>
<th>UA</th>

<th>WA</th>
<th>UA</th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN&#x002B;<break/>GRU [<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>2020</td>
<td>70.39</td>
<td>71.72</td>
<td>GM-TCN [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>2022</td>
<td>91.39</td>
<td>90.48</td>
<td>TSP&#x002B;INCA<break/> [<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>2021</td>
<td>87.43</td>
<td>87.43</td>
</tr>
<tr>
<td>SPU&#x002B;<break/>CNN [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>2021</td>
<td>66.60</td>
<td>68.40</td>
<td>LightSER<break/> [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>2022</td>
<td>94.21</td>
<td>94.15</td>
<td>GM-TCN [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>2022</td>
<td>87.35</td>
<td>87.64</td>
</tr>
<tr>
<td>LightSER<break/> [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>2022</td>
<td>70.23</td>
<td>70.76</td>
<td>CPAC [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>2023</td>
<td>94.95</td>
<td>94.22</td>
<td>CPAC [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>2023</td>
<td>89.03</td>
<td>88.41</td>
</tr>
<tr>
<td>MHA&#x002B;<break/>DRN [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>2019</td>
<td>&#x2013;</td>
<td>67.40</td>
<td>MF-CNN [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>2023</td>
<td>93.31</td>
<td>&#x2013;</td>
<td>MF-CNN [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>2023</td>
<td>94.18</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>SSL-Att [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>2023</td>
<td>&#x2013;</td>
<td>75.60</td>
<td>D-CNN [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>2023</td>
<td>95.04</td>
<td>&#x2013;</td>
<td>D-CNN [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>2023</td>
<td>95.15</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Tim-Net [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>2024</td>
<td>71.65</td>
<td>72.50</td>
<td>Tim-Net [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>2024</td>
<td>95.70</td>
<td>95.17</td>
<td>Tim-Net [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>2024</td>
<td>92.08</td>
<td>91.93</td>
</tr>
<tr>
<td>W2V [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>2024</td>
<td>75.90</td>
<td>72.10</td>
<td>Entropy [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>2024</td>
<td>&#x2013;</td>
<td>87.48</td>
<td>Entropy [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>2024</td>
<td>&#x2013;</td>
<td>79.64</td>
</tr>
<tr>
<td><bold>Our (AKD)</bold></td>
<td><bold>2025</bold></td>
<td><bold>84.45</bold></td>
<td><bold>83.34</bold></td>
<td><bold>Our (AKD)</bold></td>
<td><bold>2025</bold></td>
<td><bold>97.07</bold></td>
<td><bold>96.04</bold></td>
<td><bold>Our (AKD)</bold></td>
<td><bold>2025</bold></td>
<td><bold>97.06</bold></td>
<td><bold>95.50</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-3fn1" fn-type="other"><p>Note: The bold entries show the proposed system performace.</p></fn></table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experimental Results</title>
<sec id="s4_4_1">
<title>Ablation Study</title>
<p>This section presents additional experiments to evaluate the effectiveness of our proposed model. Various architectures are tested, including solely deep learning, combinations of neural networks, encoders/decoders, knowledge distillation, and adaptive knowledge distillation, as shown in <xref ref-type="table" rid="table-4">Table 4</xref>. According to the outcomes in <xref ref-type="table" rid="table-4">Table 4</xref>, the architecture incorporating adaptive knowledge distillation achieved the highest accuracy.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation study of the proposed model using different architectures and suggested teacher and student models Configuration on IEMOCAP corpus.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>WA</th>
<th>UA</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Teacher Model</bold></td>
<td>68.30</td>
<td>68.00</td>
</tr>
<tr>
<td>Vanilla Transformers</td>
<td>71.53</td>
<td>72.10</td>
</tr>
<tr>
<td>Hybrid Transformers Model (<bold>Student</bold>)</td>
<td>75.44</td>
<td>73.50</td>
</tr>
<tr>
<td>Teacher &#x002B; Knowledge Distillation (KD)</td>
<td>70.20</td>
<td>71.66</td>
</tr>
<tr>
<td>Teacher &#x002B; Adaptive Knowledge Distillation (AKD)</td>
<td>74.20</td>
<td>76.20</td>
</tr>
<tr>
<td>Vanilla Transformers &#x002B; Adaptive Knowledge Distillation</td>
<td>75.43</td>
<td>75.30</td>
</tr>
<tr>
<td>Multi-head Attentions &#x002B; Adaptive Knowledge Distillation</td>
<td>76.00</td>
<td>77.00</td>
</tr>
<tr>
<td><bold>Proposed Adaptive Knowledge Distillation</bold></td>
<td><bold>84.45</bold></td>
<td><bold>83.34</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-4fn1" fn-type="other"><p>Note: The bold entries show the proposed system performace.</p></fn></table-wrap-foot>
</table-wrap>
<p>The results indicate that even seemingly unrelated labels are valuable when using knowledge distillation. The knowledge distillation model serves as our initial benchmark, against which outcomes from experiments with diverse architectures are compared, as mentioned in <xref ref-type="table" rid="table-3">Table 3</xref>. Our analysis highlights the importance of knowledge exchange among non-target classes for successful logit distillation. The effectiveness of logit distillation is determined mainly by strategic adaptation, which has been understudied. Our suggested architecture combination incorporates a knowledge distillation approach to enhance the effectiveness of distillation by fine-tuning coefficients.</p>

</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Model Comparison</title>
<p>The proposed model is compared with state-of-the-art techniques using the same dataset and evaluation metrics, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. The output highlights the robustness of our method with innovative architecture in the SER domain. Our model outperformed the recent SER model and especially beat [<xref ref-type="bibr" rid="ref-39">39</xref>], knowledge distillation, and [<xref ref-type="bibr" rid="ref-38">38</xref>] multilayer attention mechanism with distillation for speech recognition, which achieved the best results recently. However, our model has a reasonable recognition rate and outperforms the recent baseline methods. These findings underscore the uniqueness and broad applicability of the features acquired through our proposed encoder-based adaptive knowledge distillation architecture. Furthermore, our model effectively captures salient information in emotion recognition tasks, as demonstrated by the confusion matrices in <xref ref-type="fig" rid="fig-2">Figs. 2</xref>&#x2013;<xref ref-type="fig" rid="fig-4">4</xref>, which provide intuitive visualizations of its performance across all evaluated datasets.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Our AKD model: confusion among actual and predicted labels of the IEMOCAP dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Our AKD model: confusion among actual and predicted labels of the EmoDB dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Our AKD model: confusion among actual and predicted labels of the RAVDESS dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79697-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Computational Analysis for Edge Devices</title>
<p>The study set out to create an <italic>AKD</italic> model tailored for resource-constrained devices, such as those on the edge. To make it lean and efficient without sacrificing accuracy, carefully applied several techniques: pruning away unnecessary parts, simplifying calculations through quantization, and using distillation to learn from larger models. This method used the ONNX standard to neatly package the model&#x2019;s parameters and weights for smooth, real-time deployment, with the results of this optimization detailed in <xref ref-type="table" rid="table-5">Table 5</xref>. Beyond that, it further boosted its speed by adapting it to lower-precision numbers (like fp16 floating-point), ensuring performance wasn&#x2019;t compromised. Ultimately, aimed to build a quick model that doesn&#x2019;t drain much power and still makes excellent predictions.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance assessment of the AKD model using various deployment frameworks on an edge device (Jetson).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Framework</th>
<th>FLOPs (G)</th>
<th>Params (M)</th>
<th>FPS (CPU)</th>
<th>FPS (GPU)</th>
<th>FPS (Jetson)</th>
<th>Model Size (MB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Keras FP32</td>
<td>80.0</td>
<td>25.5</td>
<td>15.50</td>
<td>19.00</td>
<td>11.00</td>
<td>130.00</td>
</tr>
<tr>
<td>TensorFlow FP32</td>
<td>79.8</td>
<td>25.5</td>
<td>16.50</td>
<td>20.00</td>
<td>10.50</td>
<td>130.00</td>
</tr>
<tr>
<td>PyTorch FP32</td>
<td>79.2</td>
<td>23.0</td>
<td>18.00</td>
<td>21.00</td>
<td>16.00</td>
<td>109.00</td>
</tr>
<tr>
<td>ONNX FP32</td>
<td>78.5</td>
<td>20.2</td>
<td>25.50</td>
<td>28.50</td>
<td>30.00</td>
<td>65.00</td>
</tr>
<tr>
<td>ONNX FP16</td>
<td>78.5</td>
<td>20.2</td>
<td>28.00</td>
<td>31.50</td>
<td>40.00</td>
<td>64.50</td>
</tr>
<tr>
<td>TensorRT FP32</td>
<td>77.9</td>
<td>19.8</td>
<td>33.00</td>
<td>39.90</td>
<td>45.50</td>
<td>45.00</td>
</tr>
<tr>
<td><bold>TensorRT FP16</bold></td>
<td><bold>77.9</bold></td>
<td><bold>19.8</bold></td>
<td><bold>40.00</bold></td>
<td><bold>37.00</bold></td>
<td><bold>63.00</bold></td>
<td><bold>40.00</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-5fn1" fn-type="other"><p>Note: The bold entries show the proposed system performace.</p></fn></table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Limitations of the Proposed AKD Model</title>
<p>Our model can occasionally become perplexed and incorrectly label similar or closely related emotions, such as confusing Frustration with Anger or Happiness with Excitement. Furthermore, when dealing with highly imbalanced data, our model tends to misclassify emotions as the one with the most data samples.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>The proposed system employed an adaptive knowledge distillation strategy utilizing spatio-temporal encoders/decoders in the student network, along with pre-trained Wav2Vec-2.0 (large) in the teacher networks, to enhance the model&#x2019;s performance. Our distillation model for emotion recognition leverages knowledge of non-target classes to learn discriminative features. Through our experiments on the IEMOCAP, EmoDB, and RAVDESS corpora to achieve 84.45%, 97.07%, 97.06% weighted, and 83.34%, 96.04%, 95.50% Unweighted accuracy, respectively. According to the experimental results, the student model demonstrates a strong ability to recognize emotion from speech under moderate-level noisy conditions when guided by the teacher model.</p>
<p>Furthermore, our plan is to explore the use of knowledge distillation in <italic>SER</italic> to incorporate noise into audio data, making our model more robust for real-world scenarios. Future research could enhance the proposed system by optimizing it for real-time applications, exploring various fusion techniques, addressing privacy concerns, integrating additional modalities, and evaluating the model&#x2019;s interpretability.</p>
</sec>
</body>
<back>
<ack>
<p>The authors express their appreciation and thanks to the SafeStream team for their contribution to the development of Next-Gen Multimodal AI for Improved Detection, Recognition, and Scene Analysis in UAV Applications. The authors would also like to express their gratitude to the AI-based tools used during this research to enhance it.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Ministry of Education of the Republic of Korea and the National Research Foundation of Korea (NRF-2025S1A5C3A02009153).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization, Mustaqeem Khan and Ufaq Khan; methodology, Mustaqeem Khan; software, Mustaqeem Khan and Ufaq Khan; validation, Mustaqeem Khan and Guiyoung Son; formal analysis, Mamoun Awad, Nazar Zaki and Soonil Kwon; investigation, Mustaqeem Khan, Nazar Zaki and Soonil Kwon; writing&#x2014;original draft preparation, Mustaqeem Khan and Ufaq Khan; writing&#x2014;review and editing, Mamoun Awad, Nazar Zaki, Guiyoung Son and Soonil Kwon; visualization, Mustaqeem Khan, Guiyoung Son, Nazar Zaki and Soonil Kwon; supervision, Nazar Zaki and Soonil Kwon; project administration, Guiyoung Son and Soonil Kwon; funding acquisition, Guiyoung Son and Soonil Kwon. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The study utilized publicly available datasets that can be accessed through the following links: IEMOCAP (<ext-link ext-link-type="uri" xlink:href="https://sail.usc.edu/iemocap/">https://sail.usc.edu/iemocap/</ext-link> or <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/samuelsamsudinng/iemocap-emotion-speech-database">https://www.kaggle.com/datasets/samuelsamsudinng/iemocap-emotion-speech-database</ext-link>), EmoDB (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/piyushagni5/berlin-database-of-emotional-speech-emodb">https://www.kaggle.com/datasets/piyushagni5/berlin-database-of-emotional-speech-emodb</ext-link>), and RAVDESS (<ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/uwrfkaggler/ravdess-emotional-speech-audio">https://www.kaggle.com/datasets/uwrfkaggler/ravdess-emotional-speech-audio</ext-link>).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bachate</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Suchitra</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Sentiment analysis and emotion recognition in social media: a comprehensive survey</article-title>. <source>Appl Soft Comput</source>. <year>2025</year>;<volume>174</volume>(<issue>3</issue>):<fpage>112958</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2025.112958</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Anand</surname> <given-names>RV</given-names></string-name>, <string-name><surname>Md</surname> <given-names>AQ</given-names></string-name>, <string-name><surname>Sakthivel</surname> <given-names>G</given-names></string-name>, <string-name><surname>Padmavathy</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mohan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dama&#x0161;evi&#x010D;ius</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Acoustic feature-based emotion recognition and curing using ensemble learning and CNN</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>166</volume>(<issue>4</issue>):<fpage>112151</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112151</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prabhakar</surname> <given-names>GA</given-names></string-name>, <string-name><surname>Basel</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dutta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>CVR</given-names></string-name></person-group>. <article-title>Multichannel CNN-BLSTM architecture for speech emotion recognition system by fusion of magnitude and phase spectral features using DCCA for consumer applications</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2023</year>;<volume>69</volume>(<issue>2</issue>):<fpage>226</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tce.2023.3236972</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sharma</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>A</given-names></string-name></person-group>. <article-title>DREAM: deep learning-based recognition of emotions from multiple affective modalities using consumer-grade body sensors and video cameras</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2024</year>;<volume>70</volume>(<issue>1</issue>):<fpage>1434</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lak</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Boostani</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alenizi</surname> <given-names>FA</given-names></string-name>, <string-name><surname>Mohammed</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Fakhrahmad</surname> <given-names>SM</given-names></string-name></person-group>. <article-title>RoBERTa, ResNeXt and BiLSTM with self-attention: the ultimate trio for customer sentiment analysis</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>164</volume>:<fpage>112018</fpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Basak</surname> <given-names>S</given-names></string-name>, <string-name><surname>Agrawal</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jena</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gite</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bachute</surname> <given-names>M</given-names></string-name>, <string-name><surname>Pradhan</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Challenges and limitations in speech recognition technology: a critical review of speech signal processing algorithms, tools and systems</article-title>. <source>Comput Model Eng Sci</source>. <year>2023</year>;<volume>135</volume>(<issue>2</issue>):<fpage>1053</fpage>&#x2013;<lpage>89</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2022.021755</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chauhan</surname> <given-names>GS</given-names></string-name>, <string-name><surname>Saxena</surname> <given-names>A</given-names></string-name>, <string-name><surname>Nahta</surname> <given-names>R</given-names></string-name>, <string-name><surname>Meena</surname> <given-names>YK</given-names></string-name></person-group>. <article-title>Hierarchical attention for aspect extraction using LSTM in fine-grained sentiment analysis and evaluation</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>167</volume>:<fpage>112408</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112408</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Dean</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Distilling the knowledge in a neural network</article-title>. <comment>arXiv:1503.02531. 2015</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>You</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Learning with single-teacher multi-student</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>. <publisher-loc>Menlo Park, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2018</year>. Vol. 32, p. <fpage>4390</fpage>&#x2013;<lpage>7</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nakashole</surname> <given-names>N</given-names></string-name>, <string-name><surname>Flauger</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Knowledge distillation for bilingual dictionary induction</article-title>. In: <conf-name>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2020</year>. p. <fpage>2497</fpage>&#x2013;<lpage>506</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Du</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>LF</given-names></string-name></person-group>. <article-title>Lightweight intrusion detection model based on CNN and knowledge distillation</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>165</volume>(<issue>1&#x2013;2</issue>):<fpage>112118</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2024.112118</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>ExtRe: extended temporal-spatial network for consumer-electronic WiFi-based human activity recognition</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2025</year>;<volume>71</volume>(<issue>1</issue>):<fpage>230</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tce.2024.3435881</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>G</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>D</given-names></string-name></person-group>. <article-title>EMSN: an energy-efficient memristive sequencer network for human emotion classification in mental health monitoring</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2023</year>;<volume>69</volume>(<issue>4</issue>):<fpage>1005</fpage>&#x2013;<lpage>16</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Andreas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mavromoustakis</surname> <given-names>CX</given-names></string-name>, <string-name><surname>Song</surname> <given-names>H</given-names></string-name>, <string-name><surname>Batalla</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Optimisation of CNN through transferable online knowledge for stress and sentiment classification</article-title>. <source>IEEE Trans Consum Electron</source>. <year>2024</year>;<volume>70</volume>(<issue>1</issue>):<fpage>3088</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tce.2023.3319111</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Baevski</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Auli</surname> <given-names>M</given-names></string-name></person-group>. <article-title>wav2vec 2.0: a framework for self-supervised learning of speech representations</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2020</year>;<volume>33</volume>:<fpage>12449</fpage>&#x2013;<lpage>60</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hsu</surname> <given-names>WN</given-names></string-name>, <string-name><surname>Bolte</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tsai</surname> <given-names>YHH</given-names></string-name>, <string-name><surname>Lakhotia</surname> <given-names>K</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name></person-group>. <article-title>HuBERT: self-supervised speech representation learning by masked prediction of hidden units</article-title>. <source>IEEE/ACM Trans Audio Speech Lang Process</source>. <year>2021</year>;<volume>29</volume>:<fpage>3451</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taslp.2021.3122291</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>WavLM: large-scale self-supervised pre-training for full stack speech processing</article-title>. <source>IEEE J Sel Top Signal Process</source>. <year>2022</year>;<volume>16</volume>(<issue>6</issue>):<fpage>1505</fpage>&#x2013;<lpage>18</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Woodland</surname> <given-names>PC</given-names></string-name></person-group>. <article-title>Knowledge distillation for neural transducers from large self-supervised pre-trained models</article-title>. In: <conf-name>ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>8527</fpage>&#x2013;<lpage>31</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Predicting multi-codebook vector quantization indexes for knowledge distillation</article-title>. In: <conf-name>ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kurata</surname> <given-names>G</given-names></string-name>, <string-name><surname>Saon</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Knowledge distillation from offline to streaming RNN transducer for end-to-end speech recognition</article-title>. In: <conf-name>Interspeech 2020&#x2014;The 21st Annual Conference of the International Speech Communication Association; 2020 Oct 25&#x2013;29; Shanghai, China</conf-name>. p. <fpage>2117</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Le</surname> <given-names>D</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mahadeokar</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Transformer-based acoustic modeling for hybrid speech recognition</article-title>. In: <conf-name>ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>6874</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Multilingual recurrent neural networks with residual learning for low-resource speech recognition</article-title>. In: <conf-name>INTERSPEECH 2017&#x2014;The 18th Annual Conference of the International Speech Communication Association; 2017 Aug 20&#x2013;24; Stockholm, Sweden</conf-name>. p. <fpage>704</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Al-Dujaili</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Ebrahimi-Moghadam</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Speech emotion recognition: a comprehensive survey</article-title>. <source>Wirel Pers Commun</source>. <year>2023</year>;<volume>129</volume>(<issue>4</issue>):<fpage>2525</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11277-023-10244-3</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Koren</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Factorization meets the neighborhood: a multifaceted collaborative filtering model</article-title>. In: <conf-name>Proceedings of the 14th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2008</year>. p. <fpage>426</fpage>&#x2013;<lpage>34</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Relational knowledge distillation</article-title>. In: <conf-name>Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>3967</fpage>&#x2013;<lpage>76</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Busso</surname> <given-names>C</given-names></string-name>, <string-name><surname>Bulut</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Kazemzadeh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mower</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>IEMOCAP: interactive emotional dyadic motion capture database</article-title>. <source>Lang Resour Eval</source>. <year>2008</year>;<volume>42</volume>:<fpage>335</fpage>&#x2013;<lpage>59</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Burkhardt</surname> <given-names>F</given-names></string-name>, <string-name><surname>Paeschke</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rolfes</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sendlmeier</surname> <given-names>WF</given-names></string-name>, <string-name><surname>Weiss</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mertens</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A database of German emotional speech</article-title>. In: <conf-name>INTERSPEECH 2005&#x2014;Eurospeech, 9th European Conference on Speech Communication and Technology; 2005 Sep 4&#x2013;8; Lisbon, Portugal</conf-name>. p. <fpage>1517</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Livingstone</surname> <given-names>SR</given-names></string-name>, <string-name><surname>Russo</surname> <given-names>FA</given-names></string-name></person-group>. <article-title>The Ryerson audio-visual database of emotional speech and song (RAVDESS): a dynamic, multimodal set of facial and vocal expressions in North American English</article-title>. <source>PLoS One</source>. <year>2018</year>;<volume>13</volume>(<issue>5</issue>):<fpage>e0196391</fpage>; <pub-id pub-id-type="pmid">29768426</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Silamu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A lightweight model based on separable convolution for speech emotion recognition</article-title>. In: <conf-name>INTERSPEECH 2020&#x2014;The 21st Annual Conference of the International Speech Communication Association; 2020 Oct 25&#x2013;29; Shanghai, China</conf-name>. p. <fpage>3331</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>XC</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>XZ</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>CL</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>GM-TCNet: gated multi-scale temporal convolutional network using emotion causality for speech emotion recognition</article-title>. <source>Speech Commun</source>. <year>2022</year>;<volume>145</volume>:<fpage>21</fpage>&#x2013;<lpage>35</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tuncer</surname> <given-names>T</given-names></string-name>, <string-name><surname>Dogan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Acharya</surname> <given-names>UR</given-names></string-name></person-group>. <article-title>Automated, accurate speech emotion recognition system using twine shuffle pattern and iterative neighborhood component analysis techniques</article-title>. <source>Knowl Based Syst</source>. <year>2021</year>;<volume>211</volume>:<fpage>106547</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2020.106547</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Efficient speech emotion recognition using multi-scale CNN and attention</article-title>. In: <conf-name>ICASSP 2021&#x2014;International Conference on Acoustics, Speech, and Signal Processing</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>3020</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aftab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Morsali</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ghaemmaghami</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lech</surname> <given-names>M</given-names></string-name></person-group>. <article-title>LIGHT-SERNET: a lightweight fully convolutional neural network for speech emotion recognition</article-title>. In: <conf-name>ICASSP 2022&#x2014;International Conference on Acoustics, Speech, and Signal Processing</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>6912</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wen</surname> <given-names>XC</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>J</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>XZ</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>CL</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CTL-MTNet: a novel CapsNet and transfer learning-based mixed task net for single-corpus and cross-corpus speech emotion recognition</article-title>. In: <conf-name>Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI-22)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>2305</fpage>&#x2013;<lpage>11</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Dilated residual network with multi-head self-attention for speech emotion recognition</article-title>. In: <conf-name>International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2019)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>6675</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhangale</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kothandaraman</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Speech emotion recognition based on multiple acoustic features and deep convolutional neural network</article-title>. <source>Electronics</source>. <year>2023</year>;<volume>12</volume>(<issue>4</issue>):<fpage>839</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics12040839</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kakouros</surname> <given-names>S</given-names></string-name>, <string-name><surname>Stafylakis</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mo&#x0161;ner</surname> <given-names>L</given-names></string-name>, <string-name><surname>Burget</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Speech-based emotion recognition with self-supervised models using attentive channel-wise correlations and label smoothing</article-title>. In: <conf-name>ICASSP 2023&#x2014;2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bhangale</surname> <given-names>KB</given-names></string-name>, <string-name><surname>Kothandaraman</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Speech emotion recognition using the novel PEmoNet (Parallel Emotion Network)</article-title>. <source>Appl Acoust</source>. <year>2023</year>;<volume>212</volume>(<issue>2</issue>):<fpage>109613</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apacoust.2023.109613</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>XC</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Temporal modeling matters: a novel temporal emotional modeling approach for speech emotion recognition</article-title>. In: <conf-name>ICASSP 2023&#x2014;2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>LW</given-names></string-name>, <string-name><surname>Rudnicky</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Exploring Wav2vec 2.0 fine tuning for improved speech emotion recognition</article-title>. In: <conf-name>ICASSP 2023&#x2014;2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mishra</surname> <given-names>SP</given-names></string-name>, <string-name><surname>Warule</surname> <given-names>P</given-names></string-name>, <string-name><surname>Deb</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Speech emotion recognition using MFCC-based entropy feature</article-title>. <source>Signal Image Video Process</source>. <year>2024</year>;<volume>18</volume>(<issue>1</issue>):<fpage>153</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11760-023-02716-7</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>