<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">72286</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.072286</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Multimodal Sentiment Analysis Method Based on Multi-Granularity Guided Fusion</article-title>
<alt-title alt-title-type="left-running-head">A Multimodal Sentiment Analysis Method Based on Multi-Granularity Guided Fusion</alt-title>
<alt-title alt-title-type="right-running-head">A Multimodal Sentiment Analysis Method Based on Multi-Granularity Guided Fusion</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Zilin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Liu</surname><given-names>Yan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>ms.liuyan@foxmail.com</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Jia</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Hou</surname><given-names>Senbao</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Yuping</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Chenyuan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Henan Key Laboratory of Cyberspace Situation Awareness, Key Laboratory of Cyberspace Security, Ministry of Education, Information Engineering University</institution>, <addr-line>Zhengzhou, 450001</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>State Key Laboratory of Mathematical Engineering and Advanced Computing, Information Engineering University</institution>, <addr-line>Zhengzhou, 450001</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>Henan Key Laboratory of Imaging and Intelligent Processing, Information Engineering University</institution>, <addr-line>Zhengzhou, 450001</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Yan Liu. Email: <email>ms.liuyan@foxmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year></pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>09</day><month>12</month><year>2025</year>
</pub-date>
<volume>86</volume>
<issue>2</issue>
<fpage>1</fpage>
<lpage>14</lpage>
<history>
<date date-type="received">
<day>23</day>
<month>08</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>09</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_72286.pdf"></self-uri>
<abstract>
<p>With the growing demand for more comprehensive and nuanced sentiment understanding, Multimodal Sentiment Analysis (MSA) has gained significant traction in recent years and continues to attract widespread attention in the academic community. Despite notable advances, existing approaches still face critical challenges in both information modeling and modality fusion. On one hand, many current methods rely heavily on encoders to extract global features from each modality, which limits their ability to capture latent fine-grained emotional cues within modalities. On the other hand, prevailing fusion strategies often lack mechanisms to model semantic discrepancies across modalities and to adaptively regulate modality interactions. To address these limitations, we propose a novel framework for MSA, termed Multi-Granularity Guided Fusion (MGGF). The proposed framework consists of three core components: (i) Multi-Granularity Feature Extraction Module, which simultaneously captures both global and local emotional features within each modality, and integrates them to construct richer intra-modal representations; (ii) Cross-Modal Guidance Learning Module (CMGL), which introduces a cross-modal scoring mechanism to quantify the divergence and complementarity between modalities. These scores are then used as guiding signals to enable the fusion strategy to adaptively respond to scenarios of modality agreement or conflict; (iii) Cross-Modal Fusion Module (CMF), which learns the semantic dependencies among modalities and facilitates deep-level emotional feature interaction, thereby enhancing sentiment prediction with complementary information. We evaluate MGGF on two benchmark datasets: MVSA-Single and MVSA-Multiple. Experimental results demonstrate that MGGF outperforms the current state-of-the-art model CLMLF on MVSA-Single by achieving a 2.32% improvement in F1 score. On MVSA-Multiple, it surpasses MGNNS with a 0.26% increase in accuracy. These results substantiate the effectiveness of MGGF in addressing two major limitations of existing methods&#x2014;insufficient intra-modal fine-grained sentiment modeling and inadequate cross-modal semantic fusion.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multimodal sentiment analysis</kwd>
<kwd>cross-modal fusion</kwd>
<kwd>cross-modal guided learning</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2022YFB3102904</award-id>
</award-group>
<award-group id="awg2">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>U23A20305</award-id>
<award-id>62472440</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid advancement of social media, smart devices, and multimodal sensing technologies, users are increasingly generating vast amounts of multimodal data&#x2014;such as text, images, and audio&#x2014;during daily communication. These data not only carry rich semantic content but also deeply reflect individual emotional expressions, thereby offering critical support for building more natural and emotionally intelligent human-computer interaction systems. Against this backdrop, Multimodal Sentiment Analysis (MSA) [<xref ref-type="bibr" rid="ref-1">1</xref>] has emerged as a pivotal task in the field of affective computing and has attracted growing interest across interdisciplinary domains, including Natural Language Processing (NLP), Computer Vision (CV), and Artificial Intelligence (AI). MSA aims to effectively integrate information from heterogeneous modalities to accurately recognize users&#x2019; emotional states, enhancing the system&#x2019;s capability to perceive, understand, and respond to human affect.</p>
<p>Despite the emergence of numerous MSA methods and their impressive results on various benchmark tasks, there remain two critical limitations in existing research. First, most approaches heavily rely on modeling global features of each modality while overlooking the rich local structural information embedded within modalities. For example, the method proposed by Tsai et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] focuses primarily on global image features and neglects the extraction of fine-grained emotional cues from local image regions. Similarly, although Zong et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] achieve improvements in cross-modal collaborative modeling, their method is still based on holistic modal embeddings, lacking effective representation of spatial and temporal local structures. Second, existing fusion strategies are often restricted to token-level static alignment and weighted concatenation, lacking the capability to perceive semantic discrepancies across modalities or to adaptively regulate their integration. For instance, the word-level fusion method by Chen et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] performs reasonably well in modality alignment but remains constrained by token-based concatenation strategies and fails to model local temporal dynamics and fine-grained cross-modal dependencies. Overall, most mainstream MSA methods tend to ignore intra-modal local emotional cues, which often arise from localized token combinations&#x2014;such as facial expressions or hand gestures in images&#x2014;and carry crucial emotional information. Therefore, how to effectively extract local features and achieve cooperative modeling between global and local representations remains a major challenge in multimodal sentiment analysis.</p>
<p>To address the above challenges, we propose a novel framework for multimodal sentiment analysis, termed Multi-Granularity Guided Fusion (MGGF). The framework is composed of three key modules: (i) Multi-Granularity Feature Extraction Module: This module jointly extracts both global and local emotional features from each modality and performs intra-modal fusion of multi-granularity information to construct richer unimodal representations. (ii) Cross-Modal Guidance Learning Module (CMGL): A novel mechanism based on Softmax operations and Kullback-Leibler (KL) divergence is introduced to measure the distributional divergence and complementarity between image and text modalities. This information serves as a guidance signal to inform subsequent fusion strategies, especially in contexts of modality agreement or conflict. (iii) Cross-Modal Fusion Module (CMF): We design a bidirectional interactive attention mechanism to capture fine-grained semantic dependencies between modalities. This allows for deeper emotional information exchange and contributes to the construction of more discriminative fused representations. The main contributions of this paper are summarized as follows:
<list list-type="bullet">
<list-item>
<p>We propose a multi-granularity guided fusion framework for multimodal sentiment analysis, which jointly models global and local features from both image and text modalities and achieves coordinated intra- and inter-modal information fusion.</p></list-item>
<list-item>
<p>We design a CMGL module that introduces a feature distribution divergence modeling method based on Softmax and KL divergence. This module effectively quantifies the distributional bias between modalities and guides the fusion process in both modality-consistent and modality-conflicting scenarios.</p></list-item>
<list-item>
<p>We propose a CMF with a novel bidirectional semantic interaction attention mechanism, capable of capturing fine-grained semantic dependencies across modalities and improving the sentiment discriminative power of the fused representation.</p></list-item>
<list-item>
<p>We conduct comprehensive experiments on two widely-used MSA datasets&#x2014;MVSA-Single and MVSA-Multiple. Results show that MGGF surpasses the current state-of-the-art model CLMLF on MVSA-Single by 2.32% in F1 score, and outperforms MGNNS on MVSA-Multiple with a 0.26% improvement in accuracy. These findings validate the effectiveness and potential of MGGF in multimodal sentiment analysis tasks.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work on Multimodal Sentiment Analysis</title>
<p>Early research on multimodal sentiment analysis predominantly focused on two primary fusion strategies: early fusion and late fusion. Early fusion approaches typically involve the direct concatenation of feature representations from different modalities (e.g., visual and textual) at the feature level. While such methods are relatively simple to implement and can retain a substantial amount of raw information, they often fail to model the intricate and subtle semantic correlations between modalities adequately. This limitation can lead to the inclusion of redundant or noisy features, thereby compromising the overall performance of sentiment recognition. In contrast, late fusion methods perform sentiment classification independently within each modality and subsequently integrate the outputs at the decision level.</p>
<p>Although this strategy can enhance the accuracy of intra-modal modelling to a certain extent, the lack of a collaborative cross-modal learning mechanism makes it difficult to capture deep inter-modal interactions, ultimately constraining further performance gains. To address these limitations, recent research has increasingly shifted toward intermediate fusion strategies. These approaches aim to introduce inter-modal interaction mechanisms during the process of deep feature representation learning, enabling a more fine-grained and dynamic modelling of cross-modal sentiment representations. Recent work by Wang et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] demonstrates that cross-modal hierarchical fusion with multi-task learning achieves superior performance compared with existing models on CH-SIMS, CMU-MOSI, and CMU-MOSEI datasets. Li et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] proposed a Fine-grained Multimodal Fusion Network (FMFN), which integrates learnable denoising tokens, token-level cross-modal alignment, and correlation-aware fusion to improve the performance of multimodal sentiment analysis. Yadav and Vishwakarma [<xref ref-type="bibr" rid="ref-7">7</xref>] proposed DMLANet, which integrates dual attention across image channels and spatial dimensions, alongside semantic and self-attention networks for fine-grained fusion of image-text emotional features. Zhu et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] introduced ITIN, aligning region-word pairs and using gating mechanisms to achieve fine-grained cross-modal sentiment modelling by integrating visual and textual contexts. Huang et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] developed TeFNA, a text-centric fusion framework incorporating cross-modal attention to address alignment and fusion challenges in MSA. Wang et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] proposed TETFN, which employs a text-guided multi-head attention and cross-modal mapping structure, emphasising the dominant role of text to improve both modality consistency and semantic diversity. By enabling deeper interaction between modalities, intermediate fusion achieves a balance between representation richness and alignment precision, making it a powerful alternative to early and late fusion approaches in complex sentiment analysis tasks.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Method</title>
<sec id="s3_1">
<label>3.1</label>
<title>Problem Definition</title>
<p>MSA primarily aims to comprehensively utilize the information embedded in different modalities to predict the sentiment polarity or intensity of the input samples. Formally, let us define a multimodal dataset <italic>D</italic> &#x003D; {<italic>X</italic>, <italic>Y</italic>}, where each sample <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:mi>D</mml:mi></mml:math></inline-formula> consists of multimodal input data and the corresponding sentiment label. This can be denoted as <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>U</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:math></inline-formula> represents the sequential information of the <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>m</mml:mi></mml:math></inline-formula>-th modality, and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>y</mml:mi></mml:math></inline-formula> denotes the sentiment label. In general scenarios, the modality set can include text, images, videos, or audio.</p>
<p>The core research problem addressed in this paper is how to effectively leverage the distributional differences between any two modalities to guide the fusion strategy. More specifically, for a given pair of modality sequences <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>U</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>U</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>, we first extract both global and local features, denoted as <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mi>l</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>2</mml:mn><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mn>2</mml:mn><mml:mi>l</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. These features are then fused to obtain unimodal representations <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>X</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>X</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>. Subsequently, by quantifying the distributional discrepancy between these two modality-specific representations, we derive a difference factor <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, which serves as a dynamic weighting signal to modulate the cross-modal fusion process. This mechanism allows us to preserve more unimodal-specific features when the modalities are consistent, while enhancing cross-modal interactive features in the case of conflicting modalities. Finally, the fused representation and the unimodal representations are jointly input into a classifier through a hierarchical fusion operation to predict sentiment polarity or intensity.</p>
<p>In this study, we specifically focus on the task of sentiment analysis involving text and image modalities. For ease of notation, we denote the textual sequence as <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>U</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and the image sequence as <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>U</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Overview</title>
<p>This paper proposes a MGGF, designed to enhance the modeling capability of complex cross-modal interactions in sentiment analysis tasks. MGGF strengthens adaptive perception of modality discrepancies and enhances expressive representation by guiding the fusion of global and local features from textual and visual modalities. As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the overall framework is composed of four key modules: (i) Multi-granularity Feature Extraction Module: This module employs pretrained modality encoders to obtain both global and local representations for each modality, thereby capturing comprehensive semantic information as well as fine-grained contextual dependencies. (ii) CMGL: By quantifying distributional discrepancies between unimodal representations, this module measures the heterogeneity between textual and visual modalities. The resulting cross-modal discrepancy score is then used to dynamically regulate subsequent fusion strategies. (iii) CMF: Based on the extracted multi-granularity features within each modality, this module introduces a cross-modal semantic interaction mechanism. It guides the fusion of unimodal vectors with cross-modal complementary features, thereby adapting to varying degrees of modality discrepancies and constructing a unified multimodal representation. (iv) Classifier: Finally, the fused multimodal representations and unimodal features are jointly fed into a classifier. Under the guidance of the cross-modal discrepancy score, the classifier performs the final prediction of sentiment polarity or sentiment intensity.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The overview of our proposed aspect-oriented model MGGM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_72286-fig-1.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Multi-Granularity Feature Extraction</title>
<p>The module extracts global and local semantics from text and images using pretrained Transformer encoders (BERT, ViT). As extraction is not the focus, standard models are used. Global and local features of modality <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> are denoted <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>m</mml:mi></mml:math></inline-formula> as <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup></mml:math></inline-formula>.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Global Feature Extraction</title>
<p>To capture overall semantics, we model global features by taking the [<italic>CLS</italic>] token (denoted as <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>) from the final encoder layer of modality <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>m</mml:mi></mml:math></inline-formula>, denoted <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula>, and projecting it via a feedforward layer into a unified feature space.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mi>F</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>U</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:math></inline-formula> denotes the original input sequence of modality <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>m</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Local Feature Extraction</title>
<p>To preserve fine-grained semantics, local features are extracted from all BERT tokens and ViT patches (<inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>B</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>). These embeddings are refined via Conv1D and aggregated with AdaMaxPool1d, yielding the local representations:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>t</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mn>1</mml:mn><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mn>1</mml:mn><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>B</mml:mi><mml:mi>E</mml:mi><mml:mi>R</mml:mi><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mi>X</mml:mi><mml:mi>v</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mn>1</mml:mn><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mn>1</mml:mn><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>To obtain rich modality-specific features, global and local representations of modality <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>m</mml:mi></mml:math></inline-formula> are fused via nonlinear projection to form the unimodal representation:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>O</mml:mi><mml:mi>N</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup></mml:math></inline-formula> denote the global and local feature representations of modality <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>m</mml:mi></mml:math></inline-formula>, respectively.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Intra-Modal Features Fusion</title>
<p>To enrich semantics, global <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> and local feature <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup></mml:math></inline-formula> features of modality <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>m</mml:mi></mml:math></inline-formula> are concatenated and fused via an MLP, producing the unified unimodal representation <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup></mml:math></inline-formula>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>X</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>O</mml:mi><mml:mi>N</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>X</mml:mi><mml:mi>m</mml:mi><mml:mi>l</mml:mi></mml:msubsup></mml:math></inline-formula> denote the global and local representations of modality <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>m</mml:mi></mml:math></inline-formula>, respectively, and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> corresponds to the textual and visual modalities.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Cross-Modal Guidance Learning</title>
<p>To measure representational differences between text and image, features <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> are projected into probability distributions <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>P</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>P</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> via a Softmax layer:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Softmax</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>P</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Softmax</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>The average KL divergence between <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>P</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>P</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> is used to quantify modality discrepancy, defined as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:msub><mml:mi>p</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>a</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>K</mml:mi><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo>&#x2225;</mml:mo><mml:mi>p</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>p</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>n</mml:mi></mml:math></inline-formula> denotes the sample size, and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>p</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> represent the probability distributions of visual and textual modalities, respectively.</p>
<p>To bound the discrepancy score in <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and improve learnability, the averaged KL divergence is passed through a Sigmoid function, producing the final cross-modal score <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The score <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> indicates modality divergence: lower values imply aligned representations, higher values reflect greater discrepancies. During inference, <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> serves as a dynamic weight&#x2014;emphasizing cross-modal interaction when divergence is high, and preserving unimodal features when low&#x2014;enabling adaptive semantic representation.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Cross-Modal Fusion</title>
<p>CMF captures semantic interactions between modalities, enhancing multimodal sentiment analysis, especially under modality inconsistencies. It takes <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> as input and computes a cross-modal attention matrix <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>W</mml:mi></mml:math></inline-formula> to model semantic correlations:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:math></inline-formula> denotes the feature dimension, used to scale the dot product for training stability. The <bold>softmax</bold> function ensures that the attention weights are normalized.</p>
<p>Using the weight matrices, cross-modal features are constructed to enable dynamic semantic transfer and fusion across modalities, yielding the interaction representations:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>To further capture more complex interactions between the two modalities, we apply an outer product operation between <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula>, forming the final fusion representation:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>v</mml:mi></mml:msub></mml:math></disp-formula></p>
<p>here, <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula> denotes the outer product, which explicitly encodes bilinear relationships between modalities.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Classifier</title>
<p>In the classification stage, the model constructs its input via adaptive hierarchical fusion of unimodal features&#x2014;derived from intra-modal fusion of global and local representations&#x2014;and cross-modal features generated by the CMF, which capture semantic interactions between modalities:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2295;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2295;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>X</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:math></inline-formula> represents features from the visual modality, <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>X</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> represents features from the textual modality, and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>a</mml:mi></mml:math></inline-formula> denotes the cross-modal discrepancy score computed by the cross-modal guidance module, which quantifies the degree of semantic divergence between modalities. The symbol <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>&#x2295;</mml:mo></mml:math></inline-formula> refers to hierarchical fusion operations.</p>
<p>This design adaptively balances feature contributions: high cross-modal discrepancy boosts reliance on cross-modal features to resolve conflicts, while low discrepancy favors unimodal features to reduce fusion noise. The final fused representation <italic>X</italic> is then passed to an MLP for sentiment prediction:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Optimization Object</title>
<p>We adopt a joint optimization strategy combining regression and classification losses to improve both fine-grained sensitivity and coarse-grained accuracy. The total loss is:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>here, <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> uses <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>L</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> loss (MAE) to capture subtle sentiment variations and resist outliers:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> applies cross-entropy to enhance class separability and discriminative power:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>This multi-task approach improves robustness and generalization by jointly learning fine- and coarse-grained sentiment signals.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment</title>
<sec id="s4_1">
<label>4.1</label>
<title>Set up</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Datasets</title>
<p>In our study, to ensure the validity and reliability of experimental results, we adopt two widely used benchmark datasets for MSA: MVSA-Single [<xref ref-type="bibr" rid="ref-11">11</xref>] and MVSA-Multiple. Both datasets consist of image-text pairs, with each pair annotated with three sentiment labels: Positive, Neutral, and Negative. The detailed statistics are provided in <xref ref-type="table" rid="table-1">Table 1</xref>. The two datasets differ primarily in their annotation mechanisms: MVSA-Single is annotated by a single annotator for each image-text pair, whereas MVSA-Multiple is annotated independently by three annotators, with no mutual influence among their judgments. To guarantee the consistency of sentiment annotations across modalities and ensure the effectiveness of fused representations, we performed cleaning and preprocessing on both datasets.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Sentiment label distribution in MVSA-Single and MVSA-Multiple</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Positive</th>
<th>Neutral</th>
<th>Negative</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td>MVSA-Single</td>
<td>2680</td>
<td>466</td>
<td>1365</td>
<td>4511</td>
</tr>
<tr>
<td>MVSA-Multiple</td>
<td>11,445</td>
<td>4185</td>
<td>1394</td>
<td>17,024</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Specifically, for the MVSA-Single dataset, we first removed samples with sentiment conflicts between the image and text modalities (e.g., where the image expresses positive sentiment while the text conveys negative sentiment). Next, for samples labeled as &#x201C;Neutral,&#x201D; we retained only those where the overall tendency was non-neutral, assigning the dominant sentiment as the final label. After filtering, a total of 4511 sentiment-consistent image-text pairs were preserved. For the MVSA-Multiple dataset, we followed the same procedure to remove sentiment-conflicted samples. Then, a majority voting mechanism was applied: if at least two out of three annotators agreed on the sentiment label of a sample, that label was adopted as the final annotation; otherwise, the sample was discarded. Ultimately, we obtained 17,024 high-consistency labeled image-text pairs.</p>
<p>To mitigate potential training instability caused by distributional imbalance or domain bias across datasets, we partitioned each dataset into training, validation, and testing subsets with a ratio of 8:1:1. This ensures the stability and generalization capability of the model during training, hyperparameter tuning, and evaluation. The sample distributions across subsets are reported in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Dataset split statistics for MVSA-Single and MVSA-Multiple</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Training</th>
<th>Testing</th>
<th>Validation</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td>MVSA-Single</td>
<td>3609</td>
<td>451</td>
<td>451</td>
<td>4511</td>
</tr>
<tr>
<td>MVSA-Multiple</td>
<td>13,620</td>
<td>1702</td>
<td>1702</td>
<td>17,024</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Evaluation Guidelines</title>
<p>In our experiments, to comprehensively evaluate the model&#x2019;s performance and practical applicability in multimodal sentiment analysis (MSA) tasks, we adopt two commonly used but essential evaluation metrics: (1) Accuracy (Acc): representing the overall correctness of sentiment analysis, used to measure the model&#x2019;s ability to make accurate predictions across all samples. (2) F1-score (F1): which considers both Precision and Recall in classification tasks, reflecting the balance of performance across different sentiment categories. Together, these metrics provide a well-rounded evaluation of the model&#x2019;s multimodal sentiment analysis capability. By leveraging both benchmark datasets, we aim to achieve a more precise and comprehensive understanding of the model&#x2019;s effectiveness in MSA tasks.</p>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Implementation Details</title>
<p>In this study, we employ BERT and ViT as the feature extractors for textual and visual modalities, respectively, to capture semantic information from multimodal inputs. The image input size is standardized to 224 <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224, and the hidden layer dimension is set to 768. Model training is performed using the Adam optimizer with an initial learning rate of 2e&#x2013;5. All experiments are conducted under the PyTorch 2.1.0 framework, with the runtime environment configured on Ubuntu 22.04, utilizing CUDA 12.1 and an NVIDIA GeForce RTX 4090 GPU to accelerate training. The batch size is set to 128, and the training process runs for 20 epochs. Under this configuration, the final model achieves its optimal performance.</p>
</sec>
<sec id="s4_1_4">
<label>4.1.4</label>
<title>Baseline Comparisons</title>
<p>To evaluate the effectiveness of the proposed MGGF model, we compare it with several representative and widely used baselines in multimodal sentiment analysis:
<list list-type="bullet">
<list-item>
<p>Co-MemNet [<xref ref-type="bibr" rid="ref-12">12</xref>]: A co-memory attention model using iterative memory hops to establish semantic associations&#x2014;text guides image region localisation, while images help extract sentiment keywords from text.</p></list-item>
<list-item>
<p>MVAN [<xref ref-type="bibr" rid="ref-13">13</xref>]: Proposes a multi-view attention network trained on the TumEmo dataset, extracting global, object, and scene-level image features, and jointly modelling them with text via multimodal attention.</p></list-item>
<list-item>
<p>MGNNS [<xref ref-type="bibr" rid="ref-14">14</xref>]: A graph neural network model that fuses scene and object information from images with latent textual sentiment, achieving deep-level cross-modal feature interaction and aggregation.</p></list-item>
<list-item>
<p>CLMLF [<xref ref-type="bibr" rid="ref-15">15</xref>]: Employs Transformer-based cross-layer fusion to align textual and visual features, and introduces a contrastive learning task guided by labels and data to capture sentiment-related commonalities in multimodal inputs.</p></list-item>
<list-item>
<p>MVCN [<xref ref-type="bibr" rid="ref-16">16</xref>]: proposed a Multi-View Calibration Network, which addresses the challenges of modality fusion, feature misalignment, and label inconsistency. To this end, the framework sequentially introduces a text-guided fusion module, a feature constraint task based on sentiment consistency, and an adaptive loss calibration strategy.</p></list-item>
<list-item>
<p>MFGFN [<xref ref-type="bibr" rid="ref-17">17</xref>]: proposed a Multi-Granularity Feature Gated Fusion Network that integrates fine-grained unimodal features extracted by BERT and ResNet with coarse-grained multimodal features obtained via CLIP. By introducing a co-attention encoder and a gated fusion mechanism, the model effectively enables cross-modal and multi-granularity information interaction, while adaptively assigning feature weights to enhance fusion performance.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Results</title>
<p>To assess the effectiveness of the proposed MGGF model, we compare it against several state-of-the-art multimodal sentiment analysis methods on two widely used benchmarks: MVSA-Single and MVSA-Multiple. As summarised in <xref ref-type="table" rid="table-3">Table 3</xref>, MGGF consistently outperforms baseline models, demonstrating superior multimodal sentiment discrimination. On MVSA-Single, MGGF exceeds the strongest baseline, CLMLF, by 1.25% in accuracy and 2.32% in F1-score. On MVSA-Multiple, it surpasses MGNNS and MVAN, achieving improvements of 0.63% in accuracy and 0.26% in F1-score.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance comparison of different methods on MVSA-Single and MVSA-Multiple</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th colspan="2">MVSA-Single</th>
<th colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th></th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Co-MemNet</td>
<td>70.51</td>
<td>70.01</td>
<td>69.92</td>
<td>69.83</td>
</tr>
<tr>
<td>MVAN</td>
<td>72.98</td>
<td>72.98</td>
<td>72.36</td>
<td>72.30</td>
</tr>
<tr>
<td>MGNNS</td>
<td>73.77</td>
<td>72.70</td>
<td>72.49</td>
<td>69.34</td>
</tr>
<tr>
<td>CLMLF</td>
<td>75.33</td>
<td>73.46</td>
<td>72.00</td>
<td>69.83</td>
</tr>
<tr>
<td>MFGFN</td>
<td>76.22</td>
<td>75.38</td>
<td>70.82</td>
<td>69.94</td>
</tr>
<tr>
<td><bold>MGGF (ours)</bold></td>
<td><bold>76.58</bold></td>
<td><bold>75.78</bold></td>
<td><bold>73.12</bold></td>
<td><bold>72.56</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>CLMLF integrates multi-layer features and adopts contrastive learning to enhance cross-modal alignment; however, it fails to handle modality conflicts and neglects local emotional cues, resulting in decreased accuracy. MGNNS, while based on graph neural networks, suffers from a static graph structure that lacks the capacity for dynamic adjustment and does not capture fine-grained information, leading to unstable performance. ITIN employs a region-word alignment strategy for cross-modal interaction; however, its rigid alignment mechanism introduces noise under modality inconsistency, thereby undermining recognition accuracy. DMLANet prioritizes image-side modeling while overlooking textual representations, which limits its effectiveness on text-dominant samples. TeFNA, by adopting a fixed text-centered fusion paradigm, lacks adaptability, especially in image-dominant or modality-conflicted scenarios. In contrast, the proposed MGGF combines global and local emotional features, enhancing intra-modal representational capacity. It further introduces a KL divergence-based discrepancy scoring mechanism to regulate the cross-modal fusion strategy dynamically. This design enables MGGF to effectively adapt to both modality-consistent and modality-conflicting scenarios, thereby significantly improving both accuracy and robustness, and addressing the structural limitations observed in existing approaches. These gains confirm MGGF&#x2019;s ability to effectively capture high-level semantic interactions between text and image modalities, enhancing both accuracy and generalisation.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Ablation Study</title>
<p>To further investigate the effectiveness of each component in MGGF, we conduct three sets of experiments.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Effectiveness of Each Component</title>
<p>We conducted a comprehensive ablation study on the MVSA-Single and MVSA-Multiple datasets to evaluate the individual contributions of each component within the MGGF model. The detailed experimental settings and results are summarised in <xref ref-type="table" rid="table-4">Table 4</xref>, showing performance changes as key modules are progressively removed.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation study on the architecture design of MGGF on two datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>ID</th>
<th colspan="4">Module Settings</th>
<th colspan="2">MVSA-Single</th>
<th colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th></th>
<th>Global</th>
<th>Local</th>
<th>CMGL</th>
<th>CMF</th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>&#x2713;</td>
<td></td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>76.28</td>
<td>74.83</td>
<td>72.77</td>
<td>72.06</td>
</tr>
<tr>
<td>2</td>
<td></td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>74.67</td>
<td>74.22</td>
<td>71.94</td>
<td>71.48</td>
</tr>
<tr>
<td>3</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td></td>
<td>&#x2713;</td>
<td>75.13</td>
<td>74.79</td>
<td>71.96</td>
<td>71.73</td>
</tr>
<tr>
<td>4</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td></td>
<td>74.75</td>
<td>73.96</td>
<td>72.27</td>
<td>71.82</td>
</tr>
<tr>
<td>5</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><bold>76.58</bold></td>
<td><bold>75.78</bold></td>
<td><bold>73.12</bold></td>
<td><bold>72.56</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the first set of experiments, we only utilized the global features of each modality, excluding the local features. The results show that removing local features led to performance degradation: specifically, on the MVSA-Single dataset, the metrics decreased by 0.3% and 0.95%, while on the MVSA-Multiple dataset, the drops were 0.35% and 0.5%, respectively. These findings indicate that although global features serve as the primary carriers of information, local features still provide valuable complementary cues that contribute positively to fine-grained sentiment representation.</p>
<p>In the second set of experiments, we only retained the local features of each modality. The results revealed a more severe performance drop compared to the first set of experiments. Specifically, on the MVSA-Single dataset, the performance decreased by 1.91% and 0.95%, while on the MVSA-Multiple dataset, the declines were 1.18% and 1.08%, respectively. These findings suggest that global features play a dominant role in capturing overall semantic information, whereas local features primarily serve as complementary cues to enrich fine-grained sentiment expressions.</p>
<p>In the third set of experiments, we removed the CMGL, treating unimodal and cross-modal features as equally important. The results show that, on the MVSA-Single dataset, performance dropped by 1.45% and 0.99%, while on the MVSA-Multiple dataset, it decreased by 1.16% and 0.83%. This indicates that unimodal and cross-modal features contribute differently in terms of expressive content, and the KL divergence-based mechanism for quantifying distributional discrepancies is indispensable in guiding the fusion process.</p>
<p>In the fourth set of experiments, we removed the original CMF and instead adopted a simple concatenation strategy to combine features from both modalities. The results show that, on the MVSA-Single dataset, the performance dropped by 1.8% and 1.82%, while on the MVSA-Multiple dataset, the decreases were 0.85% and 0.74%. These outcomes demonstrate that the fusion features constructed through the cross-modal interaction mechanism capture semantic dependencies and complementary relationships more effectively than simple concatenation operations.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Cross-Modal Guided Learning Analysis</title>
<p>To evaluate the impact of different distance measurement methods on model performance, we designed two variants of the MGGF framework, each employing a different similarity calculation strategy to model the relationship between textual and visual features. Specifically, MGGF-COS utilizes cosine similarity as the distance function, while MGGF-DIS is based on Euclidean distance. Both approaches estimate distances directly in the feature space without involving probability modeling. As shown in <xref ref-type="table" rid="table-5">Table 5</xref>, all three variants achieve competitive performance on multimodal sentiment analysis tasks, further validating the effectiveness of the cross-modal guidance mechanism in sentiment representation learning. Notably, MGGF-DL outperforms both MGGF-COS and MGGF-DIS. The superiority of MGGF-DL can be attributed to the use of the KL divergence mechanism, which models the discrepancies between modalities in the probability distribution space. This enables the framework to effectively capture the uncertainty and fine-grained distributional differences of cross-modal features. In contrast, although MGGF-COS and MGGF-DIS demonstrate efficiency and directness in feature space computation, their reliance solely on fixed vector similarity metrics limits their ability to model distributional information. As a result, they struggle to fully reflect the nuanced modality discrepancies underlying sentiment expressions.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance comparison of different distance measurement methods in guided learning methods</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th colspan="2">MVSA-Single</th>
<th colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th></th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>MGGF-COS</td>
<td>75.37</td>
<td>75.08</td>
<td>72.77</td>
<td>72.16</td>
</tr>
<tr>
<td>MGGF-DIS</td>
<td>74.66</td>
<td>73.94</td>
<td>72.45</td>
<td>71.92</td>
</tr>
<tr>
<td><bold>MGGF-KL</bold></td>
<td><bold>76.58</bold></td>
<td><bold>75.78</bold></td>
<td><bold>73.12</bold></td>
<td><bold>72.56</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Cross-Modal Fusion Methods</title>
<p>To assess the impact of different cross-modal fusion strategies on model performance, we designed two variants of the MGGF framework, each replacing the original cross-modal fusion module with an alternative mechanism: 1) MGGF-CAT: directly concatenates unimodal features after fusing global and local representations. 2) MGGF-CNN: introduces convolutional neural networks (CNNs) to fuse multimodal features, capturing local dependencies between modalities through receptive fields. The experimental results, presented in <xref ref-type="table" rid="table-6">Table 6</xref>, demonstrate that both alternative approaches consistently underperform compared to the original MGGF model, further confirming the effectiveness and necessity of the proposed CMF. Specifically, MGGF-CAT shows significant performance degradation, suggesting that simple concatenation of multimodal features fails to establish deep semantic interactions, thereby limiting the discriminative power of fused representations. Although MGGF-CNN leverages convolutional kernels to model local interactions between modalities, its limited receptive field restricts the capture of broader cross-domain dependencies, resulting in suboptimal fusion outcomes.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Performance comparison between different crossmodal fusion methods</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th colspan="2">MVSA-Single</th>
<th colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th></th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
<th>Acc. (%)</th>
<th>F1 (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>MGGF-CAT</td>
<td>74.75</td>
<td>73.96</td>
<td>72.27</td>
<td>71.82</td>
</tr>
<tr>
<td>MGGF-CNN</td>
<td>75.48</td>
<td>74.69</td>
<td>72.74</td>
<td>72.31</td>
</tr>
<tr>
<td><bold>MGGF</bold></td>
<td><bold>76.58</bold></td>
<td><bold>75.78</bold></td>
<td><bold>73.12</bold></td>
<td><bold>72.56</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Complexity Analysis</title>
<p>We evaluate the computational overhead and inference efficiency across its key components. The majority of the computational cost arises from the modality-specific encoders and the cross-modal interaction modules. Specifically, BERT and ViT are employed as the text and image encoders, respectively, both exhibiting a complexity of <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>L</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <italic>L</italic> denotes the input sequence length and <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>d</mml:mi></mml:math></inline-formula> represents the hidden dimension. Local feature enhancement is implemented via 1D convolution followed by adaptive pooling, incurring relatively low computational cost. The cross-modal guided learning module maps unimodal features into probability distributions and estimates their semantic discrepancy using KL divergence, with a total complexity of <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The cross-modal fusion module leverages an attention mechanism with complexity <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and introduces an outer product operation to capture high-order interactions across modalities, contributing an additional <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> complexity, making it the most computationally intensive part of the model. Under the experimental configuration, the average inference time per sample is approximately 100&#x2013;160 ms, which is acceptable for most real-time or near-real-time applications.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper proposes a Multi-Granularity Guided Fusion (MGGF) framework for multimodal sentiment analysis, aiming to improve cross-modal feature interaction and sentiment classification accuracy. By jointly constructing global and local sentiment representations within each modality and integrating them through CMGL and CMF, MGGF adaptively fuses unimodal and cross-modal features, effectively addressing misclassification caused by semantic inconsistencies across modalities. Overall, MGGF offers a robust solution to the core challenge of insufficient feature representation and ineffective fusion strategies in current multimodal sentiment analysis research.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Discussion</title>
<p>Although the proposed MGGF framework achieves competitive accuracy and robustness in multimodal sentiment analysis tasks, its overall computational overhead remains relatively high. In particular, the high-order interaction operations within the cross-modal fusion module significantly impact inference efficiency. To enhance the model&#x2019;s applicability in real-world deployment scenarios, future work will focus on model lightweighting. Specifically, we plan to explore low-rank interaction modeling, knowledge distillation, and parameter-sharing strategies, aiming to substantially reduce inference latency and resource consumption while maintaining strong performance.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported in part by the National Key Research and Development Program of China under Grant 2022YFB3102904, in part by the National Natural Science Foundation of China under Grant No. U23A20305 and No. 62472440.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Zilin Zhang: Conceptualization, Methodology, Formal Analysis, Writing&#x2014;Original Draft. Yan Liu: Conceptualization, Supervision, Writing&#x2014;Review &#x0026; Editing. Jia Liu: Software, Implementation, Data Curation, Experiments. Senbao Hou: Validation, Visualization, Writing&#x2014;Review &#x0026; Editing. Yuping Zhang: Resources, Investigation, Dataset Processing. Chenyuan Wang: Supervision, Project Administration, Funding Acquisition. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The dataset used in this study is publicly available at <ext-link ext-link-type="uri" xlink:href="https://mcrlab.net/research/mvsa-sentiment-analysis-on-multi-view-social-data/">https://mcrlab.net/research/mvsa-sentiment-analysis-on-multi-view-social-data/</ext-link> (accessed on 28 September 2025).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>U</given-names></string-name>, <string-name><surname>Abhishek</surname> <given-names>K</given-names></string-name>, <string-name><surname>Azad</surname> <given-names>HK</given-names></string-name></person-group>. <article-title>A survey of cutting-edge multimodal sentiment analysis</article-title>. <source>ACM Comput Surv</source>. <year>2024</year>;<volume>56</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3652149</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tsai</surname> <given-names>YHH</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>PY</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>LP</given-names></string-name></person-group>. <article-title>Integrating auxiliary information in self-supervised learning</article-title>. <comment>arXiv:2106.02869. 2021</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zong</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name></person-group>. <chapter-title>Acformer: an aligned and compact transformer for multimodal sentiment analysis</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Snoek</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ngo</surname> <given-names>CW</given-names></string-name>, <string-name><surname>Mei</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sebe</surname> <given-names>N</given-names></string-name></person-group>, editors. <source>Proceedings of the 31st ACM International Conference on Multimedia (MM&#x2019;23); 2023 Oct 29&#x2013;Nov 3</source>; <publisher-loc>Ottawa, ON, Canada. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2023</year>. p. <fpage>833</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3581783.3611974</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>PP</given-names></string-name>, <string-name><surname>Baltru&#x0161;aitis</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zadeh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>LP</given-names></string-name></person-group>. <chapter-title>Multimodal sentiment analysis with word-level fusion and reinforcement learning</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Schuller</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Weninger</surname> <given-names>F</given-names></string-name></person-group>, editors. <source>Proceedings of the 19th ACM International Conference on Multimodal Interaction (ICMI&#x2019;17); 2017 Nov 13&#x2013;17</source>; <publisher-loc>Glasgow, UK. New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2017</year>. p. <fpage>163</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3136755.3136801</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>LA</given-names></string-name></person-group>. <article-title>A cross modal hierarchical fusion multimodal sentiment analysis method based on multi-task learning</article-title>. <source>Inf Process Manag</source>. <year>2024</year>;<volume>61</volume>(<issue>3</issue>):<fpage>103675</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ipm.2024.103675</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>XF</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>XM</given-names></string-name></person-group>. <article-title>Learning fine-grained representation with token-level alignment for multimodal sentiment analysis</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>269</volume>(<issue>2</issue>):<fpage>126274</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.126274</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yadav</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vishwakarma</surname> <given-names>DK</given-names></string-name></person-group>. <article-title>A deep multi-level attentive network for multimodal sentiment analysis</article-title>. <source>ACM Trans Multimedia Comput Commun Appl</source>. <year>2023</year>;<volume>19</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3517139</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>JS</given-names></string-name></person-group>. <article-title>Multimodal sentiment analysis with image-text interaction network</article-title>. <source>IEEE Trans Multimedia</source>. <year>2022</year>;<volume>25</volume>:<fpage>3375</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2022.3160060</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>TeFNA: text-centered fusion network with crossmodal attention for multimodal sentiment analysis</article-title>. <source>Knowl Based Syst</source>. <year>2023</year>;<volume>269</volume>(<issue>4</issue>):<fpage>110502</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2023.110502</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>LH</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>X</given-names></string-name></person-group>. <article-title>TETFN: a text enhanced transformer fusion network for multimodal sentiment analysis</article-title>. <source>Pattern Recognit</source>. <year>2023</year>;<volume>136</volume>(<issue>2</issue>):<fpage>109259</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2022.109259</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Niu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name></person-group>. <chapter-title>Sentiment analysis on multi-view social data</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Schoeffmann</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chalupsky</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hung</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ngo</surname> <given-names>CW</given-names></string-name>, <string-name><surname>O&#x2019;Connor</surname> <given-names>NE</given-names></string-name></person-group>, editors. <source>Multimedia Modeling. Proceedings of the 22nd International Conference on Multimedia Modeling (MMM 2016); 2016 Jan 4&#x2013;6</source>; <publisher-loc>Miami, FL, USA. Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2016</year>. p. <fpage>15</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-27674-8_2</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>G</given-names></string-name></person-group>. <chapter-title>A co-memory network for multimodal sentiment analysis</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Nie</surname> <given-names>J-Y</given-names></string-name>, <string-name><surname>Baeza-Yates</surname> <given-names>R</given-names></string-name>, <string-name><surname>Croft</surname> <given-names>WB</given-names></string-name></person-group>, editors. In: <source>Proceedings of the 41st International ACM SIGIR Conference on Research &#x0026; Development in Information Retrieval; 2018 Jul 8&#x2013;12</source>; <publisher-loc>Ann Arbor, MI, USA. New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2018</year>. p. <fpage>929</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Image-text multimodal emotion classification via multi-view attentional network</article-title>. <source>IEEE Trans Multimedia</source>. <year>2020</year>;<volume>23</volume>:<fpage>4014</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2020.3035277</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name></person-group>. <chapter-title>Multimodal sentiment detection based on multi-channel graph neural networks</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Moens</surname> <given-names>M-F</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Specia</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yih</surname> <given-names>W-T</given-names></string-name></person-group>, editors. <source>Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers); 2021 Aug 1&#x2013;6</source>; <publisher-loc>Online. Stroudsburg, PA, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2021</year>. p. <fpage>328</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1162/coli_r_00312</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>T</given-names></string-name></person-group>. <article-title>CLMLF: a contrastive learning and multi-layer fusion method for multimodal sentiment detection</article-title>. <comment>arXiv:2204.05515. 2022</comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>Tackling modality heterogeneity with multi-view calibration network for multimodal sentiment detection</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Bouma</surname> <given-names>G</given-names></string-name>, <string-name><surname>Merlo</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nivre</surname> <given-names>J</given-names></string-name></person-group>, editors. <source>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2023 Jul 9&#x2013;14</source>; <publisher-loc>Toronto, ON, Canada. Stroudsburg, PA, USA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2023</year>. p. <fpage>5240</fpage>&#x2013;<lpage>52</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Multi-grained feature gating fusion network for multimodal sentiment analysis</article-title>. <source>Knowl Inf Syst</source>. <year>2025</year>;<volume>67</volume>(<issue>8</issue>):<fpage>6879</fpage>&#x2013;<lpage>905</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10115-025-02446-x</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>