<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">67812</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.067812</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>AMSA: Adaptive Multi-Channel Image Sentiment Analysis Network with Focal Loss</article-title>
<alt-title alt-title-type="left-running-head">AMSA: Adaptive Multi-Channel Image Sentiment Analysis Network with Focal Loss</alt-title>
<alt-title alt-title-type="right-running-head">AMSA: Adaptive Multi-Channel Image Sentiment Analysis Network with Focal Loss</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jin</surname><given-names>Xiaofang</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Li</surname><given-names>Yiran</given-names></name><email>hilyr7@163.com</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Yuying</given-names></name></contrib>
<aff id="aff-1"><institution>School of Information and Communication Engineering, Communication University of China</institution>, <addr-line>Beijing, 100024</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yiran Li. Email: <email>hilyr7@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>10</month><year>2025</year>
</pub-date>
<volume>85</volume>
<issue>3</issue>
<fpage>5309</fpage>
<lpage>5326</lpage>
<history>
<date date-type="received">
<day>13</day>
<month>5</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_67812.pdf"></self-uri>
<abstract>
<p>Given the importance of sentiment analysis in diverse environments, various methods are used for image sentiment analysis, including contextual sentiment analysis that utilizes character and scene relationships. However, most existing works employ character faces in conjunction with context, yet lack the capacity to analyze the emotions of characters in unconstrained environments, such as when their faces are obscured or blurred. Accordingly, this article presents the Adaptive Multi-Channel Sentiment Analysis Network (AMSA), a contextual image sentiment analysis framework, which consists of three channels: body, face, and context. AMSA employs Multi-task Cascaded Convolutional Networks (MTCNN) to detect faces within body frames; if detected, facial features are extracted and fused with body and context information for emotion recognition. If not, the model leverages body and context features alone. Meanwhile, to address class imbalance in the EMOTIC dataset, Focal Loss is introduced to improve classification performance, especially for minority emotion categories. Experimental results have shown that certain sentiment categories with lower representation in the dataset demonstrate leading classification accuracy, the AMSA yields a 2.53% increase compared with state-of-the-art methods.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Image sentiment analysis</kwd>
<kwd>adaptive multi-channel</kwd>
<kwd>class imbalance</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Affective computing [<xref ref-type="bibr" rid="ref-1">1</xref>] is an interdisciplinary field of study that focuses on developing and utilizing computer systems to understand, simulate, predict, and respond to human emotions. Sentiment analysis is a crucial aspect of affective computing, which deals with extracting and comprehending human emotional information from various multi-modal sources. Nowadays, the development of sentiment analysis has expanded from the initial text-based sentiment analysis [<xref ref-type="bibr" rid="ref-2">2</xref>] to multi-modal sentiment analysis including sound [<xref ref-type="bibr" rid="ref-3">3</xref>], video [<xref ref-type="bibr" rid="ref-4">4</xref>], facial expression [<xref ref-type="bibr" rid="ref-5">5</xref>], gesture estimation [<xref ref-type="bibr" rid="ref-6">6</xref>], and electroencephalography (EEG) [<xref ref-type="bibr" rid="ref-7">7</xref>]. Based on these, many studies have also explored bimodal [<xref ref-type="bibr" rid="ref-8">8</xref>] or multimodal [<xref ref-type="bibr" rid="ref-9">9</xref>] fusion approaches, including text-guided reconstruction modeling [<xref ref-type="bibr" rid="ref-10">10</xref>], modality uncertainty modeling [<xref ref-type="bibr" rid="ref-11">11</xref>], and semi-supervised modal contrastive learning [<xref ref-type="bibr" rid="ref-12">12</xref>].</p>
<p>One critical challenge in visual sentiment analysis is effectively modeling contextual information. While facial-body recognition achieves high accuracy in specific environment, psychological evidence shows scene context (environmental semantics, spatial attributes, concurrent activities) substantially contributes to emotion interpretation. As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, <xref ref-type="fig" rid="fig-1">Fig. 1a</xref> focuses only on the bride&#x2019;s facial and body cues, suggesting happiness and excitement. However, in <xref ref-type="fig" rid="fig-1">Fig. 1b</xref>, a falling cake indicates surprise. This example highlights the necessity of incorporating background context into visual emotion analysis.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Example pictures of scenes with effective emotional information. <bold>(a)</bold> Facial expression; <bold>(b)</bold> Facial expressions and background</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-1.tif"/>
</fig>
<p>Several methods have attempted to address this. Kosti et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] proposed a dual-branch CNN to extract body and scene features, but it did not differentiate the relative importance of different regions. Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] integrated Region Proposal Networks (RPN) to capture contextual cues but largely ignored facial information. CAER-Net [<xref ref-type="bibr" rid="ref-15">15</xref>] fused facial and scene information but struggled with occluded or missing faces, which are common in unconstrained environments. Similarly, Mittal et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] incorporated face, gait, and scene information. Wu et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] identified the problem of class imbalance in the EMOTIC dataset and proposed a hierarchical contextual emotion recognition method based on scene graphs, but achieved limited improvement.</p>
<p>Another fundamental challenge is class imbalance, where real-world datasets often contain disproportionately more negative than positive samples, or unequal distribution across categories [<xref ref-type="bibr" rid="ref-18">18</xref>]. This imbalance can cause models to overfit the majority of categories classes while neglecting minority ones, resulting in high overall accuracy but poor performance on minority categories. To address this, researchers have proposed three main strategies: data-level methods that adjust sample distributions [<xref ref-type="bibr" rid="ref-19">19</xref>], algorithm-level methods that integrate imbalance handling into model design [<xref ref-type="bibr" rid="ref-20">20</xref>], and hybrid approaches that combine both [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>To address the above limitations, this paper presents a contextual image sentiment analysis approach evaluated on the EMOTIC dataset [<xref ref-type="bibr" rid="ref-13">13</xref>]. An adaptive face detection branch is introduced, employing MTCNN [<xref ref-type="bibr" rid="ref-22">22</xref>] to detect faces within the body bounding box. In cases where multiple faces are identified, EfficientNet [<xref ref-type="bibr" rid="ref-23">23</xref>] is used to extract facial features. Features from the face, body, and contextual branches are subsequently fused for emotion recognition. To address class imbalance in the EMOTIC dataset, Focal Loss [<xref ref-type="bibr" rid="ref-24">24</xref>] is adopted as the classification loss function. Our experimental results demonstrate that using focal loss can significantly improve the classification accuracy of the model on the EMOTIC dataset.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Attention Mechanism</title>
<p>Contextual sentiment analysis requires the combination of semantic information and visual features in images for emotion classification, such as analyzing the interaction between characters and the environment. Common approaches include deep learning-based methods [<xref ref-type="bibr" rid="ref-25">25</xref>], CNNs [<xref ref-type="bibr" rid="ref-26">26</xref>], and attention mechanisms [<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>The attention mechanism is a crucial component in image classification. Spatial attention (SA) and channel attention (CA) [<xref ref-type="bibr" rid="ref-28">28</xref>] are two widely used attention mechanisms in deep learning. Their basic structures are illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. In the CA module, input features <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>F</mml:mi></mml:math></inline-formula> represent feature maps generated by convolutional operations. After undergoing Global Maximum Pooling (GMP) and Global Average Pooling (GAP), the results are fed into two shared-parameter MLP. Then outputs are multiplied and summed and then activated using the Sigmoid function to produce the final CA feature map, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>M</mml:mi><mml:mi>c</mml:mi></mml:math></inline-formula>. The SA module takes <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> as input, applies GAP and GMP along the channel dimension, concatenates the resulting 2D maps, and uses a convolutional layer followed by Sigmoid activation to produce the SA feature map, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>M</mml:mi><mml:mi>s</mml:mi></mml:math></inline-formula>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Schematic diagram of the channel attention module and spatial attention module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-2.tif"/>
</fig>
<p>The CBAM [<xref ref-type="bibr" rid="ref-29">29</xref>] module enhances CNNs by sequentially applying CA and SA, enabling the network to focus on &#x201C;what&#x201D; and &#x201C;where&#x201D; in feature maps. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the CBAM first models each channel through the CA module, and then models the spatial dimension of the feature map through the SA module. Finally, they are multiplied to obtain a final attention map. It is used to weight the original features, allowing the network to focus on important channels and spatial locations and extract more discriminative features.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Schematic diagram of CBAM attention module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-3.tif"/>
</fig>
<p>Given an intermediate feature map <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> as input, CBAM sequentially infers a 1D channel attention <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and a 2D spatial attention <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, the whole attention process can be summarized as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2297;</mml:mo><mml:mi>F</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2297;</mml:mo><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula> denotes element-by-element multiplication. <italic>C</italic> represents channels, <italic>H</italic> and <italic>W</italic> respectively represent the height and width of the feature map. During the multiplication, the attention values are copied accordingly: the channel attention values are broadcast along the spatial dimension and <italic>vice versa</italic>. <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:msup><mml:mi>F</mml:mi><mml:mi>&#x2033;</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> is the final refined output.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Face Detection and Face Feature Extraction</title>
<p>Face detection is the method of identifying and locating faces. It aims to establish the position and bounding box of faces in an image. Traditional approaches include machine learning techniques such as AdaBoost [<xref ref-type="bibr" rid="ref-30">30</xref>] and SVM [<xref ref-type="bibr" rid="ref-31">31</xref>], which rely on handcrafted features like Haar [<xref ref-type="bibr" rid="ref-32">32</xref>] and HOG [<xref ref-type="bibr" rid="ref-33">33</xref>]. More recent methods are based on deep learning, including Faster R-CNN [<xref ref-type="bibr" rid="ref-34">34</xref>], MTCNN [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>Face feature extraction involves creating a discriminative feature vector representation from a face image. Commonly used face feature extraction methods include Principal Component Analysis (PCA) [<xref ref-type="bibr" rid="ref-35">35</xref>], which reduces image dimensionality while preserving key variations for identity or emotion recognition. Local Binary Patterns (LBP) [<xref ref-type="bibr" rid="ref-36">36</xref>] capture local texture by comparing pixel neighborhoods, while more discriminative representations of face features can also be learned through the use of CNNs [<xref ref-type="bibr" rid="ref-37">37</xref>], such as VGGFace [<xref ref-type="bibr" rid="ref-38">38</xref>] and EfficientNet [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>MTCNN</title>
<p>Multi-task Cascaded Convolutional Networks (MTCNN) is a deep learning algorithm designed to detect faces and keypoints. It employs an image pyramid at the inference stage, which presents various bounding boxes for facial detection. Intersection over Union (IoU) is used in MTCNN to calculate the degree of overlap between two images. Suppose N and M are two regions, their intersection is denoted as <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>N</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>M</mml:mi></mml:math></inline-formula>, their concatenation is denoted as <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>N</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>M</mml:mi></mml:math></inline-formula>, and its IoU is calculated as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>M</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>MTCNN employs non-maximum suppression (NMS) [<xref ref-type="bibr" rid="ref-39">39</xref>] to eliminate redundant bounding boxes by retaining only the most confident detections. In the formula <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the score of each box. If a box <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> has an IoU with the highest-scoring box <italic>M</italic> greater than the threshold <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, its score is set to zero and it is discarded.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mi>O</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mi>O</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>MTCNN uses image coordinate back-calculation to figure out the position and size of the real frame based on the predicted values output by the network. The suggestion box is obtained as in <xref ref-type="fig" rid="fig-4">Fig. 4a</xref>, which is the blue box in <xref ref-type="fig" rid="fig-4">Fig. 4b</xref>, has the following coordinates:</p>

<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>x</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac><mml:mspace width="1em"></mml:mspace><mml:mi>x</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>y</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>h</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac><mml:mspace width="1em"></mml:mspace><mml:mi>y</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>h</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Graphical representation of image coordinate back-calculation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-4.tif"/>
</fig>
<p>Based on the position of the proposed box and the offset value of the network output, the true value can be found, which is the yellow box in <xref ref-type="fig" rid="fig-4">Fig. 4b</xref>, whose coordinates are calculated as follows:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mspace width="1em"></mml:mspace><mml:mi>x</mml:mi><mml:mn>2</mml:mn><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>y</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mi>y</mml:mi><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>y</mml:mi><mml:mn>1</mml:mn><mml:mspace width="1em"></mml:mspace><mml:mi>y</mml:mi><mml:mn>2</mml:mn><mml:mo>=</mml:mo><mml:mi>y</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>y</mml:mi><mml:mn>2</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>In the formula above, <italic>index</italic> is the index value in the graph, <italic>strides</italic> is the step size of the slide and <italic>scale</italic> is the scaling ratio in the image pyramid, <italic>w</italic>, <italic>h</italic> are the width and height of the suggestion box and <italic>offset</italic> is the offset value predicted by the network.</p>
<p>MTCNN performs face detection and keypoint localization through an image pyramid and a cascade of three CNNs: P-Net, R-Net, and O-Net. P-Net performs feature extraction and classification through a sliding window to quickly identify face candidates and their locations. R-Net leverages deeper features to classify and regress the candidate regions, resulting in more accurate face bounding boxes. O-Net further refines the bounding boxes and precisely localizes facial keypoints. As shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, the MTCNN effectively achieves the tasks of face detection and face keypoint localization using three cascaded networks that gradually screen and refine results.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Complete flow of MTCNN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-5.tif"/>
</fig>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>EfficientNet</title>
<p>EfficientNet achieves a balance of performance and efficiency through composite scaling that adjusts the width, depth, and resolution of the network. The model is pre-trained on large-scale image datasets to learn rich image features. These features can be applied to face feature extraction, accelerating model training and improving performance on face datasets. <xref ref-type="fig" rid="fig-6">Fig. 6a</xref> shows the baseline EfficientNet structure, while <xref ref-type="fig" rid="fig-6">Fig. 6b</xref>&#x2013;<xref ref-type="fig" rid="fig-6">d</xref> expand it in width, depth, and resolution. <xref ref-type="fig" rid="fig-6">Fig. 6e</xref> illustrates the main idea: scaling all three dimensions together.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>EfficientNet network structure. (<bold>a</bold>) baseline; (<bold>b</bold>) width scaling; (<bold>c</bold>) depth scaling; (<bold>d</bold>) resolution scaling; (<bold>e</bold>) compound scaling</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-6.tif"/>
</fig>
<p>Our target is to maximize the model accuracy for any given resource constraints, which can be formulated as an optimization problem:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:munder><mml:mo form="prefix">max</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi></mml:mrow></mml:munder><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mspace width="1em"></mml:mspace><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x22EF;</mml:mo><mml:mi>s</mml:mi></mml:mrow></mml:munder><mml:msubsup><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mover><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>H</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>W</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>C</mml:mi><mml:mo stretchy="false">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>M</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>y</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd></mml:mtd><mml:mtd><mml:mi>F</mml:mi><mml:mi>L</mml:mi><mml:mi>O</mml:mi><mml:mi>P</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>f</mml:mi><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Among these variables, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>r</mml:mi></mml:math></inline-formula> denote depth, width, and resolution multiplicity of the network. The subscripts <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>1</mml:mn><mml:mo>&#x2026;</mml:mo><mml:mi>s</mml:mi></mml:math></inline-formula> indicate the serial number of the stage, and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the convolution operation on the first layer of the convolution operation. <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> means that <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> has <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> convolutional layers with the same structure in the <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>i</mml:mi></mml:math></inline-formula>-th stage. Memory represents the memory usage of the model. FLOPs are used as a constraint to balance model complexity and performance. By controlling the number of FLOPs, more efficient models can be designed with limited computational resources.</p>
<p>EfficientNet proposes a compound scaling method in which a mixing factor <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>&#x03D5;</mml:mi></mml:math></inline-formula> is used to unify the scaling depth, width and resolution parameter, which is calculated as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>Scale</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mrow><mml:mtext>depth</mml:mtext></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mrow><mml:mtext>width</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mrow><mml:mtext>resolution</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Loss Function</title>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>Loss Function Used for Classification Task&#x2014;Focal Loss</title>
<p>Focal Loss [<xref ref-type="bibr" rid="ref-24">24</xref>] is designed to address class imbalance, a common challenge in machine learning where some categories have significantly fewer samples. Unlike Cross-Entropy Loss [<xref ref-type="bibr" rid="ref-40">40</xref>], which treats all classes equally, Focal Loss reweights the loss to reduce the impact of well-classified or majority-class samples, thereby focusing more on difficult and minority-class examples.</p>
<p>The focal loss for binary classification is derived from the Cross-Entropy (CE) loss.
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="1em"></mml:mspace><mml:mspace width="1em"></mml:mspace><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="1em"></mml:mspace><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mo>&#x00B1;</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> specifies the ground truth class, the <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>p</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> is the estimated probability of the model for the class labeled <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Define <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mspace width="1em"></mml:mspace><mml:mspace width="thinmathspace"></mml:mspace><mml:mspace width="thinmathspace"></mml:mspace><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:mspace width="thinmathspace"></mml:mspace><mml:mspace width="thinmathspace"></mml:mspace><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Rewritten CE:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>A remarkable property of this loss is that even easily categorized samples can produce losses of non-trivial magnitude. A common way to address class imbalances is to introduce a weighting factor <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, the CE loss of the <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> equilibrium is written as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This loss is a extension of the CE, which we consider to be the experimental baseline for focal loss. Adding a modulation factor <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> to the cross entropy loss with a tunable focusing paramete <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> balanced focal loss is used. In summary, define the focal loss as:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>F</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The focal loss method operates by scaling the loss using an exponential function that incorporates the actual observed category probability to adjust the loss weight. For instances that are easily classified, the focus factor approaches 0, which in turn limits their impact on the overall loss. For samples that are difficult to categorize, they are given a larger focus factor, which increases their loss weight and influence in training. The introduction of a focus factor allows Focal Loss to effectively decrease the impact of the majority category and concentrate on the minority category, thereby enhancing the accuracy of predictions for the latter. In our implementation, we set the hyperparameters of Focal Loss to &#x03B1; &#x003D; 0.5 and &#x03B3; &#x003D; 2, following standard practice in prior works, which offer a stable performance in our preliminary ablation experiments.</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Loss Function Used for Regression Task&#x2014;MSE</title>
<p>Mean Squared Error (MSE) Loss is commonly used in regression tasks [<xref ref-type="bibr" rid="ref-41">41</xref>] to measure the difference between predicted and true values. MSE is calculated by adding up the squares of the differences between them for each sample and dividing by the number of samples. The mathematical expression is given below:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>MSE yields a non-negative output, and values closer to 0 reflect more accurate predictions made by the model. Since the square of the error is used in the MSE calculation, relatively large errors are magnified, making the model more sensitive to large errors. It is a continuously derivable loss function, which makes it possible to apply optimization algorithms such as gradient descent for parameter updating during the training process. Computation entails only the summation of squared differences and addition, resulting in low computational complexity.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Our Approach</title>
<sec id="s3_1">
<label>3.1</label>
<title>Context Feature Extraction</title>
<p>When extracting emotional features from a scene, the entire image is fed into the channel. To capture crucial regions of the image, the CBAM attention mechanism is employed, which integrates channel and spatial attention mechanisms.</p>
<p>In this study, we choose ResNet18 [<xref ref-type="bibr" rid="ref-42">42</xref>] as the backbone network for both the body and context feature extraction channels. Compared with deeper models such as ResNet50 and ResNet101, ResNet18 achieves better performance while significantly reducing training time and computational resource usage. Its relatively simple architecture also makes it well-suited for integration with the CBAM attention mechanism.</p>
<p>Incorporating the CBAM attention mechanism on the network structure of ResNet18 can enhance the feature representation capability of the model. By adaptively learning the importance of channels and space, CBAM can provide a more accurate and robust feature representation, enabling the model to better capture the features that are most useful for the task. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows the exact location of the CBAM module when it is integrated in a ResBlock. We apply CBAM to the convolutional output in each block.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Exact location of ResBlock addition to the CBAM attention module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-7.tif"/>
</fig>
<p>The complete network structure of context feature extraction channel after adding CBAM is shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the input of the channel is the original image that contains both characters and scenes. The CBAM attention module is added in all four stages of ResNet18, and finally the features of the scene channel are obtained through the average pooling layer and the fully connected layer.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Schematic diagram of the ResNet18 network incorporating the CBAM attention module</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-8.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Face Feature Extraction</title>
<p>Since more than 25% of the images in the EMOTIC dataset fail to detect clear faces, an adaptive face channel is introduced to perform face detection and feature extraction directly from body frames. As shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, MTCNN is used to detect faces within the body bounding boxes. If one or more faces are detected, the extracted facial frames are input into EfficientNet for feature extraction. The features from the scene, body, and face channels are then fused for subsequent classification. If no face is detected, the face channel is deactivated (weight set to 0), and classification proceeds using only the body and context features.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>The structure of face feature extraction channel</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-9.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Model Structural Design</title>
<p>The complete model structure of adaptive multi-channel image sentiment analysis network with Focal Loss (AMSA) is shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Complete model structure of AMSA</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-10.tif"/>
</fig>
<p>Both the body feature extraction module and the context feature extraction module use pre-trained ResNet18 models. Additionally, the context feature extraction part incorporates the CBAM attention mechanism to extract scene-relevant features. The face feature extraction part takes the body box as input. When a face is successfully detected by MTCNN, the corresponding facial region is extracted and processed by EfficientNet, and the features from all three channels (face, body, and context) are fused for classification. If no face is detected, the face channel is deactivated (the weight is set to zero), and the model switches to a dual-branch fusion mode, using only body and context features for emotion prediction. To ensure stable performance under both conditions, we adopt fixed fusion weights determined through validation set optimization. The fusion output <italic>y</italic> is computed as follows:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0.4</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.6</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0.5</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.2</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.3</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Finally, the fusion network module combines the features from the three feature extraction modules and carries out fine-grained sentiment representation regression using two fully connected layers. The output results in both discrete and continuous dimensions.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>EMOTIC Dataset</title>
<p>The EMOTIC dataset is a collection of images of people in unconstrained environments, where over 25% of people&#x2019;s faces are partially occluded or have very low resolution. It contains 23,571 images and 34,320 annotated characters. The dataset uses a combination of queries containing keywords for different locations, social environments, diverse activities, and various emotional states, and it provides manually annotated body regions. The annotations comprise 3 continuous emotion dimensions: Valence, Arousal, and Dominance and 26 discrete emotions: anticipation, engagement, confidence, happiness, surprise, fatigue, embarrassment, anger and more. The detailed definitions of these categories can be found in [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>], and <xref ref-type="fig" rid="fig-11">Fig. 11</xref> shows example images of these categories.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Example of discrete category annotated image from EMOTIC dataset [<xref ref-type="bibr" rid="ref-13">13</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-11.tif"/>
</fig>
<p>In our experiment, the dataset is split into three subsets: 70% for training, 20% for validation, and 10% for testing. The split ensures a balanced distribution of emotion categories across the subsets.</p>
<p>Images are annotated according to the VAD model [<xref ref-type="bibr" rid="ref-44">44</xref>], which represents emotions through a combination of three continuous dimensions, each of which is an integer value within the range of 1 to 10. The VAD model is defined as follows: valence (V) is used to measure the degree of positivity or pleasantness of an emotion, ranging from negative to positive; arousal (A) is used to measure an individual&#x2019;s degree of arousal, ranging from inactive/calm to agitated/ready for action; and dominance (D) is used to measure a person&#x2019;s degree of control over a situation, ranging from non-control to control. <xref ref-type="fig" rid="fig-12">Fig. 12</xref> illustrates the various values assigned to each dimension.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Example of continuous dimension annotated image from EMOTIC dataset [<xref ref-type="bibr" rid="ref-13">13</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-12.tif"/>
</fig>
<p>We utilize average precision (AP) to measure the performance of discrete category classification and the Jaccard coefficient (JC) [<xref ref-type="bibr" rid="ref-45">45</xref>] to test the similarity between two sets. Due to the classification of EMOTIC dataset belongs to multi-label classification, the final output may have some overlap with Ground Truth or vary completely. Therefore, we use JC to represents the ratio of the size of the intersection of A and B to the size of the concatenation of A and B. It is defined by <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>, the higher the JC, the better the result. The maximum value of the JC is 1, where the detected category and the Ground Truth category are identical.
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>J</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>B</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Performance of VAD in three continuous dimensions is measured through the MAE. A lower calculated value indicates a smaller error and better predictive performance of the model as defined in <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref>.
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The distribution of the 26 emotion categories in the EMOTIC dataset is presented in <xref ref-type="fig" rid="fig-13">Fig. 13</xref>, with a descending order from left to right.</p>
<fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Percentage of each category in the dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67812-fig-13.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Analysis of Results</title>
<p>To evaluate the contribution of each module in our model, we conducted an ablation study involving various combinations of CBAM, the adaptive face channel, and Focal Loss in six experimental configurations. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, the results demonstrate that each module contributes positively to the overall performance. (iii) and (iv) demonstrate the benefit of incorporating adaptive face recognition channel and confirm that the focal loss effectively enhanced performance on imbalanced data. And the full AMSA model achieves the highest accuracy, while configurations with only partial components show reduced performance.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Ablation experiment</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>CBAM</th>
<th>Face channel</th>
<th>Focal loss</th>
<th>AP</th>
</tr>
</thead>
<tbody>
<tr>
<td>(i)</td>
<td>&#x00D7;</td>
<td>&#x00D7;</td>
<td>&#x00D7;</td>
<td>27.34</td>
</tr>
<tr>
<td>(ii)</td>
<td>&#x221A;</td>
<td>&#x00D7;</td>
<td>&#x00D7;</td>
<td>28.37</td>
</tr>
<tr>
<td>(iii)</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td>&#x00D7;</td>
<td>30.37</td>
</tr>
<tr>
<td>(iv)</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td><bold>35.04</bold></td>
</tr>
<tr>
<td>(v)</td>
<td>&#x221A;</td>
<td>&#x00D7;</td>
<td>&#x221A;</td>
<td>33.25</td>
</tr>
<tr>
<td>(vi)</td>
<td>&#x00D7;</td>
<td>&#x00D7;</td>
<td>&#x221A;</td>
<td>32.01</td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-1fn1" fn-type="other"><p>Note: Best values are shown in bold.</p></fn></table-wrap-foot>
</table-wrap>
<p>To evaluate our model on a fine-grained dimension, we conducted classification experiments on 26 discrete emotion categories and compared our AMSA model with several representative methods, including those proposed by Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>], Lee et al. [<xref ref-type="bibr" rid="ref-15">15</xref>], Mittal et al. [<xref ref-type="bibr" rid="ref-16">16</xref>], Kosti et al. [<xref ref-type="bibr" rid="ref-43">43</xref>], Wang et al. [<xref ref-type="bibr" rid="ref-46">46</xref>], Li et al. [<xref ref-type="bibr" rid="ref-47">47</xref>] and Yang et al. [<xref ref-type="bibr" rid="ref-48">48</xref>].</p>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, AMSA achieved the highest mean Average Precision (mAP), which was 2.53% higher than that of the best baseline paper [<xref ref-type="bibr" rid="ref-48">48</xref>]. This significant improvement demonstrates the effectiveness of the proposed adaptive multi-channel framework. Additionally, our model achieved the highest classification accuracy in 10 emotion categories, including Anger, Annoyance, and others. These gains can be attributed to the introduction of Focal Loss, which addresses the class imbalance issue in the EMOTIC dataset by giving more weight to underrepresented classes during training. This is especially important given the unbalanced distribution of emotion categories in the dataset. Moreover, we observed that in complex scene categories with subtle emotional cues, such as Annoyance and Sympathy, our model achieved 24.22% and 35.30%, respectively, significantly outperforming other methods. This confirms that our model does not overly rely on a single modality, but is capable of adapting to diverse visual cues present in complex scenes.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of AP on discrete dimensions</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th align="center">Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</th>
<th align="center">Lee et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</th>
<th align="center">Mittal et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</th>
<th align="center">Kosti et al. [<xref ref-type="bibr" rid="ref-43">43</xref>]</th>
<th align="center">Wang et al. [<xref ref-type="bibr" rid="ref-46">46</xref>]</th>
<th align="center">Li et al. [<xref ref-type="bibr" rid="ref-47">47</xref>]</th>
<th align="center">Yang et al. [<xref ref-type="bibr" rid="ref-48">48</xref>]</th>
<th align="center">AMSA (Ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td>(1) Affection</td>
<td><bold>46.89</bold></td>
<td>23.25</td>
<td>41.83</td>
<td>27.85</td>
<td>32.27</td>
<td>37.93</td>
<td>35.51</td>
<td>42.89</td>
</tr>
<tr>
<td>(2) Anger</td>
<td>10.87</td>
<td>9.71</td>
<td>11.41</td>
<td>9.49</td>
<td>13.04</td>
<td>13.73</td>
<td>14.6</td>
<td><bold>16.96</bold></td>
</tr>
<tr>
<td>(3) Annoyance</td>
<td>11.27</td>
<td>13.43</td>
<td>17.37</td>
<td>14.06</td>
<td>18.92</td>
<td>28.14</td>
<td>16.94</td>
<td><bold>24.22</bold></td>
</tr>
<tr>
<td>(4) Anticipation</td>
<td>62.64</td>
<td>54.12</td>
<td>67.59</td>
<td>58.64</td>
<td>56.32</td>
<td>61.08</td>
<td>89.05</td>
<td><bold>95.14</bold></td>
</tr>
<tr>
<td>(5) Aversion</td>
<td>5.93</td>
<td>8.64</td>
<td>11.71</td>
<td>7.48</td>
<td>9.82</td>
<td>9.61</td>
<td>16.83</td>
<td><bold>17.51</bold></td>
</tr>
<tr>
<td>(6) Confidence</td>
<td>72.49</td>
<td>72.35</td>
<td>65.27</td>
<td>78.35</td>
<td>75.21</td>
<td><bold>80.08</bold></td>
<td>73.11</td>
<td>78.16</td>
</tr>
<tr>
<td>(7) Disapproval</td>
<td>11.28</td>
<td>15.26</td>
<td>17.35</td>
<td>14.97</td>
<td>20.11</td>
<td>21.54</td>
<td><bold>27.45</bold></td>
<td>27.44</td>
</tr>
<tr>
<td>(8) Disconnection</td>
<td>26.91</td>
<td>21.53</td>
<td><bold>41.46</bold></td>
<td>21.32</td>
<td>30.45</td>
<td>28.32</td>
<td>31.7</td>
<td>37.2</td>
</tr>
<tr>
<td>(9) Disquietment</td>
<td>16.94</td>
<td>16.81</td>
<td>12.69</td>
<td>16.89</td>
<td>19.58</td>
<td>22.57</td>
<td>23.37</td>
<td>21.19</td>
</tr>
<tr>
<td>(10) Doubt/Confusion</td>
<td>18.68</td>
<td><bold>37.77</bold></td>
<td>31.28</td>
<td>29.63</td>
<td>21.26</td>
<td>33.50</td>
<td>19.55</td>
<td>26.24</td>
</tr>
<tr>
<td>(11) Embarrassment</td>
<td>1.94</td>
<td>2.29</td>
<td><bold>10.51</bold></td>
<td>3.18</td>
<td>2.32</td>
<td>4.16</td>
<td>7.24</td>
<td>6.27</td>
</tr>
<tr>
<td>(12) Engagement</td>
<td>88.56</td>
<td>83.43</td>
<td>84.62</td>
<td>87.53</td>
<td>87.01</td>
<td>88.12</td>
<td>94.38</td>
<td><bold>98.36</bold></td>
</tr>
<tr>
<td>(13) Esteem</td>
<td>13.33</td>
<td>17.84</td>
<td>18.79</td>
<td>17.73</td>
<td>15.09</td>
<td>20.50</td>
<td>23.01</td>
<td><bold>25.28</bold></td>
</tr>
<tr>
<td>(14) Excitement</td>
<td>71.89</td>
<td>70.68</td>
<td><bold>80</bold>.<bold>54</bold></td>
<td>77.16</td>
<td>68.64</td>
<td>80.11</td>
<td>60.42</td>
<td>79.52</td>
</tr>
<tr>
<td>(15) Fatigue</td>
<td>13.26</td>
<td>8.91</td>
<td>11.95</td>
<td>9.70</td>
<td>14.11</td>
<td><bold>17.51</bold></td>
<td>14.67</td>
<td>12.18</td>
</tr>
<tr>
<td>(16) Fear</td>
<td>4.21</td>
<td>12.36</td>
<td><bold>21.36</bold></td>
<td>14.14</td>
<td>8.68</td>
<td>15.56</td>
<td>11.23</td>
<td>10.29</td>
</tr>
<tr>
<td>(17) Happiness</td>
<td>73.26</td>
<td>55.79</td>
<td>69.51</td>
<td>58.26</td>
<td>77.59</td>
<td>76.01</td>
<td>84.24</td>
<td><bold>84.49</bold></td>
</tr>
<tr>
<td>(18) Pain</td>
<td>6.52</td>
<td>9.22</td>
<td>9.56</td>
<td>8.94</td>
<td>11.58</td>
<td>14.56</td>
<td>16.44</td>
<td><bold>19.98</bold></td>
</tr>
<tr>
<td>(19) Peace</td>
<td><bold>32.85</bold></td>
<td>19.03</td>
<td>30.72</td>
<td>21.56</td>
<td>26.18</td>
<td>26.76</td>
<td>26.05</td>
<td>21.89</td>
</tr>
<tr>
<td>(20) Pleasure</td>
<td>57.46</td>
<td>43.22</td>
<td><bold>61.89</bold></td>
<td>45.46</td>
<td>49.48</td>
<td>55.64</td>
<td>50.92</td>
<td>54.37</td>
</tr>
<tr>
<td>(21) Sadness</td>
<td>25.42</td>
<td>10.39</td>
<td>19.74</td>
<td>19.66</td>
<td><bold>39.43</bold></td>
<td><bold>30</bold>.<bold>80</bold></td>
<td>37.43</td>
<td>24.16</td>
</tr>
<tr>
<td>(22) Sensitivity</td>
<td>5.99</td>
<td>7.34</td>
<td>4.11</td>
<td>9.28</td>
<td><bold>11.34</bold></td>
<td><bold>9.59</bold></td>
<td>10.7</td>
<td>7.46</td>
</tr>
<tr>
<td>(23) Suffering</td>
<td>23.39</td>
<td>9.71</td>
<td>20.92</td>
<td>18.84</td>
<td><bold>42.35</bold></td>
<td><bold>30.70</bold></td>
<td>30.85</td>
<td>20.71</td>
</tr>
<tr>
<td>(24) Surprise</td>
<td>9.02</td>
<td>13.70</td>
<td>16.45</td>
<td><bold>18.81</bold></td>
<td>7.75</td>
<td>17.92</td>
<td>7.21</td>
<td>12.11</td>
</tr>
<tr>
<td>(25) Sympathy</td>
<td>17.53</td>
<td>16.29</td>
<td>30.68</td>
<td>14.71</td>
<td>12.28</td>
<td>15.26</td>
<td>13.66</td>
<td><bold>35</bold>.<bold>30</bold></td>
</tr>
<tr>
<td>(26) Yearning</td>
<td>10.55</td>
<td>9.89</td>
<td>10.83</td>
<td>8.29</td>
<td>8.59</td>
<td>11.01</td>
<td>8.63</td>
<td><bold>11.62</bold></td>
</tr>
<tr>
<td>mAP</td>
<td>28.42</td>
<td>25.1</td>
<td>31.53</td>
<td>27.38</td>
<td>30.17</td>
<td>32.41</td>
<td>32.51</td>
<td><bold>35.04</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-2fn1" fn-type="other"><p>Note: Best values are shown in bold.</p></fn></table-wrap-foot>
</table-wrap>
<p>In addition to the aforementioned models, several studies have reported only the mAP without disclosing the classification accuracy for individual emotion categories. For completeness, we provide a comparative summary of these results in <xref ref-type="table" rid="table-3">Table 3</xref>. de Lima Costa et al. [<xref ref-type="bibr" rid="ref-49">49</xref>] proposed a lightweight single-stream model focused on computational efficiency, but showed limited accuracy improvement. Etesam et al. [<xref ref-type="bibr" rid="ref-50">50</xref>] adopted a large language model for caption-based reasoning. Zhang et al. [<xref ref-type="bibr" rid="ref-51">51</xref>] proposed a training paradigm that extracts visual affective cues from communication using unedited data and topic-aware contextual encoding. In contrast, our AMSA reaches the highest mAP of 35.04%, wihch shows a balanced and accurate recognition result across both common and rare classes.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of mAP on discrete dimensions</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>de Lima Costa et al. [<xref ref-type="bibr" rid="ref-49">49</xref>]</th>
<th>Etesam et al. [<xref ref-type="bibr" rid="ref-50">50</xref>]</th>
<th>Zhang et al. [<xref ref-type="bibr" rid="ref-51">51</xref>]</th>
<th>AMSA (Ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td>mAP</td>
<td>30.02</td>
<td>32.55</td>
<td>32.91</td>
<td><bold>35.04</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-3fn1" fn-type="other"><p>Note: Best values are shown in bold.</p></fn></table-wrap-foot>
</table-wrap>
<p>To evaluate our model in the continuous emotional dimension, we conducted regression analysis on the VAD space. As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, our model achieves a 10% lower average error rate compared to [<xref ref-type="bibr" rid="ref-43">43</xref>], and performs comparably to [<xref ref-type="bibr" rid="ref-14">14</xref>]. This result proves the effectiveness of using Mean Squared Error (MSE) loss in our regression setting. Compared to both approaches, AMSA adopts a more comprehensive image analysis framework by incorporating adaptive face detection and multi-channel feature fusion. Unlike [<xref ref-type="bibr" rid="ref-43">43</xref>], which uses SL1 Loss, our method leverages MSE Loss to better optimize continuous predictions. While Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] also adopt MSE, AMSA achieves similar overall performance in average error but exhibits more balanced results across the three VAD dimensions.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Comparison of average error rate on continuous dimensions</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Kosti et al. [<xref ref-type="bibr" rid="ref-43">43</xref>]</th>
<th>Zhang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>]</th>
<th>AMSA (Ours)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Violence</td>
<td>0.9</td>
<td>0.7</td>
<td>0.8</td>
</tr>
<tr>
<td>Arousal</td>
<td>1.3</td>
<td>1.0</td>
<td>1.1</td>
</tr>
<tr>
<td>Dominance</td>
<td>0.9</td>
<td>1.0</td>
<td>1.0</td>
</tr>
<tr>
<td>Average</td>
<td>1.0</td>
<td>0.9</td>
<td>0.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To provide a more intuitive result for our experiment, we displays three representative samples from the EMOTIC dataset in <xref ref-type="table" rid="table-5">Table 5</xref>, including two successful cases and one challenging case with facial occlusion. The classification predictions are denoted in black for agreeing with those in Ground Truth, and in red for misclassification. We evaluated the results&#x2019; efficacy using JC coefficients and MAE. Our model correctly detects the majority of ground-truth categories in the first two examples. Compared to [<xref ref-type="bibr" rid="ref-43">43</xref>], the JC coefficients are improved by 0.15 and 0.30, respectively, while the MAE values are reduced by 1.23 and 1.57. In the third case, the face was obscured and model relied solely on body and context features for judgment, so it is not very accurate in terms of emotions, resulting in misjudgments such as fear and pain. Although its performance has declined, the gap compared to [<xref ref-type="bibr" rid="ref-43">43</xref>] remains small. This indicates the validity of the AMSA model presented in this paper.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Visualization of inference results</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col/>
<col/>
<col align="center"/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Image</th>
<th>Ground truth</th>
<th>AMSA (Ours)</th>
<th>[<xref ref-type="bibr" rid="ref-43">43</xref>]</th>
<th align="center">Ground truth</th>
<th>AMSA (Ours)</th>
<th>[<xref ref-type="bibr" rid="ref-43">43</xref>]</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-graphic mime-subtype="tif" xlink:href="CMC_67812-inline-1.tif"/></td>
<td>Disconnection Doubt/Confusion</td>
<td><styled-content style-type="color" style="color: #FF0000;">Anticipation</styled-content><break/>Disconnection<break/>Doubt/Confusion<break/><styled-content style-type="color" style="color: #FF0000;">Engagement</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Peace</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Sensitivity</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Yearning</styled-content><break/>JC: 0.28</td>
<td><styled-content style-type="color" style="color: #FF0000;">Anticipation</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Confidence</styled-content><break/>Disconnection<break/><styled-content style-type="color" style="color: #FF0000;">Peace</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Sensitivity</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Yearning</styled-content><break/>JC: 0.13</td>
<td>V: 5<break/>A: 3<break/>D: 9</td>
<td>V: 5.5<break/>A: 2.7<break/>D: 7.8<break/>MAE: 0.67</td>
<td>V: 6.0<break/>A: 3.0<break/>D: 6.7<break/>MAE: 1.90</td>
</tr>

<tr>
<td><inline-graphic mime-subtype="tif" xlink:href="CMC_67812-inline-2.tif"/></td>
<td>Affection<break/>Confidence<break/>Esteem<break/>Happiness<break/>Peace<break/>Pleasure<break/>Sympathy</td>
<td>Affection<break/><styled-content style-type="color" style="color: #FF0000;">Anticipation</styled-content><break/>Confidence<break/><styled-content style-type="color" style="color: #FF0000;">Engagement</styled-content><break/>Esteem<break/>Happiness<break/>Pleasure<break/>Sympathy<break/>JC: 0.75</td>
<td>Affection<break/><styled-content style-type="color" style="color: #FF0000;">Anticipation</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Engagement</styled-content><break/>Esteem<break/>Happiness<break/>Pleasure<break/><styled-content style-type="color" style="color: #FF0000;">Sensitivity</styled-content><break/>Sympathy<break/><styled-content style-type="color" style="color: #FF0000;">Yearning</styled-content><break/>JC: 0.45</td>
<td>V: 8<break/>A: 8<break/>D: 7.0</td>
<td>V: 7.8<break/>A: 8.2<break/>D: 7.0<break/>MAE: 0.13</td>
<td>V: 6.0<break/>A: 8.0<break/>D: 6.7<break/>MAE: 1.70</td>
</tr>
<tr>
<td><inline-graphic mime-subtype="tif" xlink:href="CMC_67812-inline-3.tif"/></td>
<td>Annoyance Anticipation Disconnection Disquietment Doubt<break/>/Confusion<break/>Sadness</td>
<td>Disconnection<break/><styled-content style-type="color" style="color: #FF0000;">Pain</styled-content><break/> Doubt<break/>/Confusion<break/><styled-content style-type="color" style="color: #FF0000;">Fear</styled-content><break/><styled-content style-type="color" style="color: #FF0000;">Embarrassment</styled-content><break/>JC: 0.33</td>
<td>Sadness<break/>Annoyance<break/><styled-content style-type="color" style="color: #FF0000;">Fatigue</styled-content><break/> Disconnection<break/><styled-content style-type="color" style="color: #FF0000;">Sensitivity</styled-content><break/>JC: 0.5</td>
<td>V: 4<break/>A: 4<break/>D: 4</td>
<td>V: 4.2<break/>A: 3.7<break/>D: 3.8<break/>MAE: 0.23</td>
<td>V: 3.9<break/>A: 4.3<break/>D: 4.1<break/>MAE: 0.17</td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-5fn1" fn-type="other"><p>Note: Categories marked in red indicate misclassified predictions.</p></fn></table-wrap-foot>
</table-wrap>
<p>This example also shows the importance of our adaptive strategy. When the face is blocked, the model automatically switches to rely more on body and context features. Although this fallback approach cannot fully replace facial information, it still enables the model make reasonable predictions by using other visual cues in the image.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, we proposed AMSA, a multi-channel model for contextual sentiment analysis that adaptively integrates facial, body, and scene features. The face channel leverages MTCNN and EfficientNet, while the body and context channels employ ResNet18, with CBAM attention further enhancing scene-level feature extraction. To address class imbalance in the EMOTIC dataset, we adopt Focal Loss for classification and MSE Loss for regression. Experimental results demonstrate that AMSA outperforms previous methods by 2.63% in average accuracy, and our model achieves leading performance on several underrepresented emotion categories, confirming the effectiveness of the proposed architecture and loss function.</p>
<p>In the future, this work can be extended to applications such as intelligent surveillance, human-computer interaction, supporting the practical deployment of image-based sentiment analysis. Further research will focus on improving the model&#x2019;s generalization in cross-domain scenarios and exploring the integration of textual and audio modalities for more comprehensive multimodal sentiment analysis.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Xiaofang Jin; methodology, Yiran Li; software, Yiran Li; validation, Yuying Yang; data curation, Yuying Yang; writing&#x2014;original draft preparation, Yuying Yang, Yiran Li; writing&#x2014;review and editing, Xiaofang Jin, Yiran Li, Yuying Yang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are available within the article [<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Picard</surname> <given-names>RW</given-names></string-name></person-group>. <source>Affective computing</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>; <year>1995</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>SentiDiff: combining textual information and sentiment diffusion patterns for twitter sentiment analysis</article-title>. <source>IEEE Trans Knowl Data Eng</source>. <year>2020</year>;<volume>32</volume>(<issue>10</issue>):<fpage>2026</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tkde.2019.2913641</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Er</surname> <given-names>MB</given-names></string-name></person-group>. <article-title>A novel approach for classification of speech emotions based on deep and acoustic features</article-title>. <source>IEEE Access</source>. <year>2020</year>;<volume>8</volume>:<fpage>221640</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2020.3043201</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Context-aware attention network for human emotion recognition in video</article-title>. <source>Adv Multimedia</source>. <year>2020</year>;<volume>2020</volume>:<fpage>8843413</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2020/8843413</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jayaraman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mahendran</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An improved facial expression recognition using CNN-BiLSTM with attention mechanism</article-title>. <source>Int J Adv Comput Sci</source>. <year>2024</year>;<volume>15</volume>(<issue>5</issue>):<fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.14569/IJACSA.2024.01505132</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Slogrove</surname> <given-names>K</given-names></string-name>, <string-name><surname>van der Haar</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Group emotion recognition in the wild using pose estimation and LSTM neural networks</article-title>. In: <conf-name>Proceedings of the International Conference on Artificial Intelligence, Big Data, Computing and Data Communication Systems (ICABCD); 2022 Aug 4&#x2013;5</conf-name>; <publisher-loc>Durban, South Africa</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>3DCANN: a spatio-temporal convolution attention neural network for EEG emotion recognition</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2022</year>;<volume>26</volume>(<issue>11</issue>):<fpage>5321</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2021.3083525</pub-id>; <pub-id pub-id-type="pmid">34033551</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Braytee</surname> <given-names>A</given-names></string-name>, <string-name><surname>Anaissi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>L</given-names></string-name>, <string-name><surname>Akram</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Ensemble pretrained models for multimodal sentiment analysis using textual and video data fusion</article-title>. In: <conf-name>Proceedings of the WWW&#x0027;24: Companion Proceedings of the ACM Web Conference 2024; 2024 May 13&#x2013;17</conf-name>; <publisher-loc>Singapore</publisher-loc>. p. <fpage>1841</fpage>&#x2013;<lpage>48</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>He</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content</article-title>. In: <conf-name>Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI); 2025 Feb 25&#x2013;Mar 4</conf-name>; <publisher-loc>Philadelphia, PA, USA</publisher-loc>. p. <fpage>1601</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nakagawa</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Text-guided reconstruction network for sentiment analysis with uncertain missing modalities</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2025</year>:<fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taffc.2025.3541743</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Trustworthy multimodal fusion for sentiment analysis in ordinal sentiment space</article-title>. <source>IEEE Trans Circuits Syst Video Technol</source>. <year>2024</year>;<volume>34</volume>(<issue>8</issue>):<fpage>7657</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tcsvt.2024.3376564</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name></person-group>. <article-title>CorMulT: a semi-supervised modality correlation-aware multimodal transformer for sentiment analysis</article-title>. <comment>arXiv:2407.07046. 2024</comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kosti</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alvarez</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Recasens</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lapedriza</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Emotion recognition in context</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>1960</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Context-aware affective graph reasoning for emotion recognition</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Multimedia and Expo (ICME 2019); 2019 Jul 8&#x2013;12</conf-name>; <publisher-loc>Shanghai, China</publisher-loc>. p. <fpage>151</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sohn</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Context-aware emotion recognition networks</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2019 Oct 27&#x2013;Nov 2</conf-name>; <publisher-loc>Seoul, Republic of Korea</publisher-loc>. p. <fpage>10142</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mittal</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guhan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bhattacharya</surname> <given-names>U</given-names></string-name>, <string-name><surname>Chandra</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bera</surname> <given-names>A</given-names></string-name>, <string-name><surname>Manocha</surname> <given-names>D</given-names></string-name></person-group>. <article-title>EmotiCon: context-aware multimodal emotion recognition using Frege&#x2019;s principle</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>14234</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Hierarchical context-based emotion recognition with scene graphs</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2024</year>;<volume>35</volume>(<issue>3</issue>):<fpage>3725</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2022.3196831</pub-id>; <pub-id pub-id-type="pmid">36018874</pub-id></mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>EA</given-names></string-name></person-group>. <article-title>Learning from imbalanced data</article-title>. <source>IEEE Trans Knowl Data Eng</source>. <year>2009</year>;<volume>21</volume>(<issue>9</issue>):<fpage>1263</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tkde.2008.239</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Grouped SMOTE with noise filtering mechanism for classifying imbalanced data</article-title>. <source>IEEE Access</source>. <year>2019</year>;<volume>7</volume>:<fpage>170668</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2019.2955086</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>An imbalance compensation framework for background subtraction</article-title>. <source>IEEE Trans Multimedia</source>. <year>2017</year>;<volume>19</volume>(<issue>11</issue>):<fpage>2425</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2017.2701645</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bader-El-Den</surname> <given-names>M</given-names></string-name>, <string-name><surname>Teitei</surname> <given-names>E</given-names></string-name>, <string-name><surname>Perry</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Biased random forest for dealing with the class imbalance problem</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2019</year>;<volume>30</volume>(<issue>7</issue>):<fpage>2163</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tnnls.2018.2878400</pub-id>; <pub-id pub-id-type="pmid">30475733</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Joint face detection and alignment using multitask cascaded convolutional networks</article-title>. <source>IEEE Signal Process Lett</source>. <year>2016</year>;<volume>23</volume>(<issue>10</issue>):<fpage>1499</fpage>&#x2013;<lpage>503</lpage>. doi:<pub-id pub-id-type="doi">10.1109/lsp.2016.2603342</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Le</surname> <given-names>QV</given-names></string-name></person-group>. <article-title>EfficientNet: rethinking model scaling for convolutional neural networks. arXiv:1905.11946. 2020</article-title>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Focal loss for dense object detection</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>;<volume>42</volume>(<issue>2</issue>):<fpage>318</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2018.2858826</pub-id>; <pub-id pub-id-type="pmid">30040631</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>J</given-names></string-name></person-group>. <article-title>SOLVER: scene-object interrelated visual emotion reasoning network</article-title>. <source>IEEE Trans Image Process</source>. <year>2021</year>;<volume>30</volume>:<fpage>8686</fpage>&#x2013;<lpage>701</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tip.2021.3118983</pub-id>; <pub-id pub-id-type="pmid">34665725</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pikoulis</surname> <given-names>I</given-names></string-name>, <string-name><surname>Filntisis</surname> <given-names>PP</given-names></string-name>, <string-name><surname>Maragos</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Leveraging semantic scene characteristics and multi-stream convolutional architectures in a contextual approach for video-based visual emotion recognition in the wild</article-title>. In: <conf-name>Proceedings of the 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG); 2021 May 15&#x2013;18</conf-name>; <publisher-loc>Jodhpur, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Context based vision emotion recognition in the wild</article-title>. In: <conf-name>Proceedings of the 17th IEEE Conference on Industrial Electronics and Applications (ICIEA); 2022 Dec 16&#x2013;19</conf-name>; <publisher-loc>Chengdu, China</publisher-loc>. p. <fpage>479</fpage>&#x2013;<lpage>84</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Albanie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Squeeze-and-excitation networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>;<volume>42</volume>(<issue>8</issue>):<fpage>2011</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2019.2913372</pub-id>; <pub-id pub-id-type="pmid">31034408</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Woo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name></person-group>. <chapter-title>CBAM: convolutional block attention module</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Ferrari</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hebert</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sminchisescu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Weiss</surname> <given-names>Y</given-names></string-name></person-group>, editors. <source>Computer vision&#x2014;ECCV 2018</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2018</year>. p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_1</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Freund</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Schapire</surname> <given-names>RE</given-names></string-name></person-group>. <article-title>A decision-theoretic generalization of on-line learning and an application to boosting</article-title>. <source>J Comput Syst Sci</source>. <year>1997</year>;<volume>55</volume>(<issue>1</issue>):<fpage>119</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1006/jcss.1997.1504</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cortes</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vapnik</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Support-vector networks</article-title>. <source>Mach Learn</source>. <year>1995</year>;<volume>20</volume>(<issue>3</issue>):<fpage>273</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1023/a:1022627411411</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Viola</surname> <given-names>PA</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>MJ</given-names></string-name></person-group>. <article-title>Rapid object detection using a boosted cascade of simple features</article-title>. In: <conf-name>Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR 2001; 2001 Dec 8&#x2013;14</conf-name>; <publisher-loc>Kauai, HI, USA</publisher-loc>. </mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jabade</surname> <given-names>V</given-names></string-name>, <string-name><surname>Ingale</surname> <given-names>A</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Robust face detection and identification using HOG-based features and machine learning</article-title>. In: <conf-name>Proceedings of the 2023 2nd International Conference on Futuristic Technologies (INCOFT); 2023 Nov 2&#x2013;3</conf-name>; <publisher-loc>Coimbatore, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Faster R-CNN: towards real-time object detection with region proposal networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2017</year>;<volume>39</volume>(<issue>6</issue>):<fpage>1137</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2016.2577031</pub-id>; <pub-id pub-id-type="pmid">27295650</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Turk</surname> <given-names>M</given-names></string-name>, <string-name><surname>Pentland</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Eigenfaces for recognition</article-title>. <source>J Cogn Neurosci</source>. <year>1991</year>;<volume>3</volume>(<issue>1</issue>):<fpage>71</fpage>&#x2013;<lpage>86</lpage>; <pub-id pub-id-type="pmid">23964806</pub-id></mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Triggs</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Enhanced local texture feature sets for face recognition under difficult lighting conditions</article-title>. <source>IEEE Trans Image Process</source>. <year>2010</year>;<volume>19</volume>(<issue>6</issue>):<fpage>1635</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tip.2010.2042645</pub-id>; <pub-id pub-id-type="pmid">20172829</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Umer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rout</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Tiwari</surname> <given-names>S</given-names></string-name>, <string-name><surname>AlZubi</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Alanazi</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Yurii</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Human-computer interaction using deep fusion model-based facial expression recognition system</article-title>. <source>Comput Model Eng Sci</source>. <year>2023</year>;<volume>135</volume>(<issue>2</issue>):<fpage>1165</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2022.023312</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Parkhi</surname> <given-names>OM</given-names></string-name>, <string-name><surname>Vedaldi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Deep face recognition</article-title>. In: <conf-name>Proceedings of the British Machine Vision Conference (BMVC); 2015 Sep 7&#x2013;10</conf-name>; <publisher-loc>Swansea, UK</publisher-loc>. </mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Felzenszwalb</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>McAllester</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ramanan</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Object detection with discriminatively trained part-based models</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2010</year>;<volume>32</volume>(<issue>9</issue>):<fpage>1627</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2009.167</pub-id>; <pub-id pub-id-type="pmid">20634557</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sangari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sethares</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Convergence analysis of two loss functions in soft-max regression</article-title>. <source>IEEE Trans Signal Process</source>. <year>2016</year>;<volume>64</volume>(<issue>5</issue>):<fpage>1280</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsp.2015.2504348</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Pyramid scene parsing network</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jun 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>6230</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kosti</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alvarez</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Recasens</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lapedriza</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Context based emotion recognition using EMOTIC dataset</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>;<volume>42</volume>(<issue>11</issue>):<fpage>2755</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2019.2916866</pub-id>; <pub-id pub-id-type="pmid">31095475</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Russell</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>A circumplex model of affect</article-title>. <source>J Pers Soc Psychol</source>. <year>1980</year>;<volume>39</volume>(<issue>6</issue>):<fpage>1161</fpage>&#x2013;<lpage>78</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jaccard</surname> <given-names>P</given-names></string-name></person-group>. <article-title>The distribution of the flora in the alpine zone</article-title>. <source>New Phytol</source>. <year>1902</year>;<volume>2</volume>:<fpage>37</fpage>&#x2013;<lpage>50</lpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Context-dependent emotion recognition</article-title>. <source>J Vis Commun Image Represent</source>. <year>2022</year>;<volume>89</volume>:<fpage>103679</fpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Human emotion recognition with relational region-level analysis</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2023</year>;<volume>14</volume>(<issue>1</issue>):<fpage>650</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taffc.2021.3064918</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Robust emotion recognition in context debiasing</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>12447</fpage>&#x2013;<lpage>57</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>de Lima Costa</surname> <given-names>W</given-names></string-name>, <string-name><surname>Talavera</surname> <given-names>E</given-names></string-name>, <string-name><surname>Figueiredo</surname> <given-names>LS</given-names></string-name>, <string-name><surname>Teichrieb</surname> <given-names>V</given-names></string-name></person-group>. <article-title>High-level context representation for emotion recognition in images</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>. <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>326</fpage>&#x2013;<lpage>34</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Etesam</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yal&#x00E7;&#x0131;n</surname> <given-names>&#x00D6;N</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lim</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Contextual emotion recognition using large vision language models</article-title>. In: <conf-name>Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2024 Oct 14&#x2013;18</conf-name>; <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>. p. <fpage>4769</fpage>&#x2013;<lpage>76</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>JZ</given-names></string-name></person-group>. <article-title>Learning emotion representations from verbal and nonverbal communication</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>18993</fpage>&#x2013;<lpage>9004</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>