<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81626</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081626</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Multi-Branch Cross-Modal Cross-Attention for Image&#x2013;Text Multimodal Sentiment Classification</article-title>
<alt-title alt-title-type="left-running-head">Multi-Branch Cross-Modal Cross-Attention for Image&#x2013;Text Multimodal Sentiment Classification</alt-title>
<alt-title alt-title-type="right-running-head">Multi-Branch Cross-Modal Cross-Attention for Image&#x2013;Text Multimodal Sentiment Classification</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Huang</surname><given-names>Xinshan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Pei</surname><given-names>Zirui</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Tan</surname><given-names>Chaohong</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Meng</surname><given-names>Zuqiang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>zqmeng@126.com</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Computer, Electronics and Information, Guangxi University</institution>, <addr-line>Nanning</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Guangxi Key Laboratory of Digital Infrastructure, Guangxi Zhuang Autonomous Region Information Center</institution>, <addr-line>Nanning</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Zuqiang Meng. Email: <email>zqmeng@126.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>90</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81626.pdf"></self-uri>
<abstract>
<p>Multimodal Sentiment Analysis (MSA) plays an important role in understanding social media content; however, existing methods often struggle with the heterogeneity and complex interactions between images and text. These challenges include inter-modal information asymmetry, insufficient feature fusion, and noise interference, which collectively limit robustness and accuracy. To address these issues, we propose a multimodal sentiment classification model termed Multi-Branch Cross-Modal Cross-Attention Gating (MB-CMCAG). The model first incorporates a Transformer-based image caption generation module to convert raw images into semantically rich auxiliary textual descriptions, which complement the original text and form paired textual inputs with enhanced visual semantics. To capture multi-source features, MB-CMCAG adopts a dual-branch feature extraction architecture: the visual branch encodes images using a Vision Transformer (ViT), while the textual branch encodes text with Bidirectional Encoder Representations from Transformers (BERT); a Contrastive Language-Image Pre-training (CLIP) model is also introduced for joint image&#x2013;text feature extraction. To exploit cross-modal correlations and enable hierarchical fusion, we construct a cross-modal attention module that supports bidirectional information flow from image to text and from text to image. Building on this, a cross-modal gated mechanism is introduced to selectively regulate the transmission and aggregation of features from different sources, thereby improving noise suppression and sentiment sensitivity. Experimental results on the public MVSA-Single and MVSA-Multiple datasets show that MB-CMCAG achieves accuracies of 76.38% and 73.87%, respectively, outperforming existing baselines by a clear margin in image&#x2013;text multimodal sentiment classification.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multimodal sentiment analysis</kwd>
<kwd>cross-modal cross-attention</kwd>
<kwd>cross-modal gated fusion</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62266004</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Open Fund of the Key Laboratory of Digital Infrastructure in Guangxi</funding-source>
<award-id>GXDINBC202401</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Sentiment analysis, as an important research direction in natural language processing, aims to identify, interpret, and process the emotions and affective states expressed by humans. With the rapid development of Internet technologies and social platforms, an increasing number of people, both domestically and internationally, use social media such as Weibo, Twitter, and Facebook to express their thoughts and feelings. At the same time, the carriers of expressed content have evolved from the traditional text modality to multiple modalities, including images, video, and audio. Emotion recognition is a crucial area of artificial intelligence, and multimodal emotion recognition has become a research hotspot in recent years [<xref ref-type="bibr" rid="ref-1">1</xref>]. Multimodal Sentiment Analysis (MSA) has broad applications across many domains, including human&#x2013;computer interaction [<xref ref-type="bibr" rid="ref-2">2</xref>] and advertising and commerce [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Existing MSA methods can be broadly categorized by their feature-fusion strategies: early fusion, intermediate fusion, and late fusion [<xref ref-type="bibr" rid="ref-4">4</xref>]. Early fusion combines modalities at the feature-extraction stage but struggles with fine-grained cross-modal information and suffers from feature redundancy [<xref ref-type="bibr" rid="ref-5">5</xref>]. Late fusion trains separate classifiers for each modality but fails to model inter-modal correlations effectively [<xref ref-type="bibr" rid="ref-6">6</xref>]. Intermediate fusion, which facilitates feature interaction within neural network layers, has gained attention for its ability to capture cross-modal relationships [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>With the development of deep learning, MSA has shifted from traditional fusion methods to more sophisticated fusion mechanisms. A growing number of researchers have begun to leverage vision&#x2013;language models such as CLIP [<xref ref-type="bibr" rid="ref-9">9</xref>] for MSA. For example, reference [<xref ref-type="bibr" rid="ref-10">10</xref>] proposed a CLIP-based sentiment analyzer that employs a high-level feature extraction module and integrates features using attention mechanisms and custom layers, thereby addressing the limitations of traditional methods in effectively combining multimodal data. In addition, reference [<xref ref-type="bibr" rid="ref-8">8</xref>] introduced a multimodal sentiment classification model based on a Gated Attention mechanism, which uses a pretrained convolutional neural network to extract fine-grained visual features and fuses them with textual features through gated attention, in order to reduce noise interference and highlight the key textual components that influence sentiment polarity.</p>
<p>Although substantial research efforts have been devoted to advancing MSA, existing methods still suffer from three major limitations: (1) reliance on a single feature extraction pathway, which hinders the full extraction of discriminative representations from images and text across multiple perspectives and scales; (2) relatively shallow cross-modal interaction mechanisms, which typically rely on simple concatenation or summation and fail to precisely capture the complementary relationships and accurate alignment between visual and textual modalities in fine-grained semantic space; and (3) static and rigid modality fusion strategies that lack dynamic screening and adaptive re-weighting of cross-modal information. In particular, many CLIP-based or interaction-based methods emphasize either global alignment or shallow cross-modal fusion, but they do not simultaneously model fine-grained local semantics, coarse-grained global consistency, and adaptive modality translation in a unified framework. Consequently, these methods are susceptible to redundant noise interference, thereby limiting sentiment discrimination performance in complex real-world scenarios.</p>
<p>To address these challenges, this paper proposes a multimodal sentiment analysis model based on a multi-branch cross-modal cross-attention gated network (MB-CMCAG). To overcome the limitations of single-path feature extraction, we introduce a dual-branch complementary encoding mechanism. The fine-grained branch captures local semantics by employing the Vision Transformer (ViT) [<xref ref-type="bibr" rid="ref-11">11</xref>] model to encode images and the Bidirectional Encoder Representations from Transformers (BERT) [<xref ref-type="bibr" rid="ref-12">12</xref>] model to encode text. The coarse-grained branch utilizes CLIP to capture global alignment. These two branches form mutually complementary representations at different semantic levels, providing a richer foundation for subsequent fusion.</p>
<p>Furthermore, to bridge the semantic gap between visual and textual modalities and enable effective dynamic fusion of semantic information, we adopt a cross-modal cross-attention module and a cross-modal gated mechanism. In addition, to enhance the model&#x2019;s understanding of visual emotional semantics, we employ a transformer-based caption generation module that converts images into captions. This transforms visual semantic information into textual form, thereby constructing an auxiliary sentence that enriches the original text. In summary, the main contributions of our work are as follows:<list list-type="bullet">
<list-item>
<p><bold>Multi-Perspective Hierarchical Feature Extraction Mechanism</bold>&#x2014;We propose a dual-branch encoding architecture that integrates fine-grained modeling (ViT&#x002B;BERT) with coarse-grained modeling (CLIP). Unlike previous single-path or dual-stream methods that operate at a single semantic scale, our framework is the first to explicitly construct complementary multi-scale representations in parallel, simultaneously capturing local details and global semantics within each modality.</p></list-item>
<list-item>
<p><bold>Hierarchical Cross-Modal Interaction and Fusion Framework</bold>&#x2014;Bidirectional cross-attention is adopted within each branch to explicitly model fine-grained semantic alignment, while a gated mechanism is introduced between branches to enable dynamic re-weighting of cross-scale information. This design goes beyond shallow and unidirectional interactions in existing methods by enabling bidirectional, multi-level semantic alignment in a true cross-modal space.</p></list-item>
<list-item>
<p><bold>Modality Translation Enhancement Strategy</bold>&#x2014;We introduce an image captioning module to explicitly convert visual semantics into textual representations and incorporate them into the fusion process. By reinforcing visual emotional signals in the textual dimension, this approach fundamentally bridges the modality gap and improves the comparability and fusibility of visual information in semantic space.</p></list-item>
<list-item>
<p><bold>Comprehensive Evaluation and Component Analysis</bold>&#x2014;Extensive experiments, ablation studies, and efficiency analysis demonstrate the effectiveness and practicality of the proposed framework.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Text Sentiment Analysis</title>
<p>Text sentiment analysis is one of the earliest and most mature research directions in affective computing and natural language processing. Early studies mainly relied on sentiment lexicons and traditional machine learning methods, whereas the recent rise of deep learning models has significantly advanced the performance of this field.</p>
<p>First, lexicon-based methods [<xref ref-type="bibr" rid="ref-13">13</xref>] primarily rely on precompiled sentiment lexicons to determine overall sentiment by identifying sentiment-bearing words and their intensities in text. For example, Alwan et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] proposed a hybrid model based on rough set theory and sentiment lexicons for polarity analysis of Arabic political articles. To address low classification accuracy on Arabic political texts, they constructed a dedicated corpus of 206 annotated documents and performed classification via three core steps&#x2014;preprocessing, feature extraction, and hybrid-model construction&#x2014;achieving an accuracy of 85.483%. In a comment-classification task, Hamouda et al. employed the SentiWordNet lexicon [<xref ref-type="bibr" rid="ref-15">15</xref>], which assigns positive, neutral, and negative sentiment scores to lexical items [<xref ref-type="bibr" rid="ref-16">16</xref>]. Although lexicon-based methods offer interpretability, require no training data, and are grounded in manual rules, they struggle to capture contextual shifts, sarcasm or metaphor, and domain-specific expressions. Traditional machine-learning approaches commonly use algorithms such as k-nearest neighbors (k-NN), support vector machines (SVM) [<xref ref-type="bibr" rid="ref-17">17</xref>], and naive Bayes (NB) [<xref ref-type="bibr" rid="ref-18">18</xref>]. Goel et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] presented a real-time Twitter sentiment-analysis system that combines naive Bayes with SentiWordNet, improving classification accuracy by incorporating SentiWordNet&#x2019;s scoring scheme. Rathor et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] compared SVM, NB, and maximum entropy (ME) for sentiment analysis of Amazon product reviews and found SVM achieved the highest accuracy. More recently, the rise of deep learning has substantially advanced text sentiment analysis. Kim [<xref ref-type="bibr" rid="ref-21">21</xref>] applied convolutional neural networks (CNNs) to sentence-level sentiment classification with strong results. Zhou et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a bilingual LSTM model with an attention mechanism for cross-lingual sentiment classification. Wang et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] introduced a tree-structured region-dividing CNN&#x2013;LSTM hybrid model for dimensional sentiment analysis; by partitioning text into multiple semantic regions, extracting local features via CNNs, and capturing long-range cross-region dependencies with LSTMs, the model produces continuous-valued predictions of sentiment intensity.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Image Sentiment Analysis</title>
<p>Image sentiment analysis aims to identify and interpret the emotions conveyed in images by analyzing their visual content. Similar to text sentiment analysis, image sentiment analysis has also evolved from traditional approaches to deep learning&#x2013;based methods.</p>
<p>Early approaches typically relied on hand-crafted visual features, such as low-level cues including color, texture, shape, and illumination, combined with machine learning classifiers for sentiment prediction. For example, Siersdorfer et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] investigated the relationship between the affective content of images in social media and their metadata and visual content. By analyzing more than 586,000 images from Flickr, they explored the feasibility of predicting image sentiment polarity by combining visual features (e.g., color distributions and SIFT descriptors) with textual sentiment labels extracted using the SentiWordNet lexicon. In [<xref ref-type="bibr" rid="ref-25">25</xref>], a large-scale Visual Sentiment Ontology (VSO) containing more than 3000 adjective&#x2013;noun pairs (ANPs) was constructed, and a SentiBank library with 1200 ANP detectors was developed on this basis, which significantly improved the accuracy of sentiment prediction for visual content by detecting affective concepts in images. Yang et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed a graph-based sentiment analysis method that jointly models user-posted images and comments from friends to infer users&#x2019; affective states. However, these methods often struggle to capture complex high-level semantic and affective information in images. With the advent of deep learning, image sentiment analysis has achieved substantial progress. Song et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a visual-attention-based image sentiment analysis method (SentiNet-A), which integrates convolutional neural networks (CNNs) with a visual attention mechanism to improve image sentiment classification performance. Islam and Zhang [<xref ref-type="bibr" rid="ref-28">28</xref>] adopted transfer learning by initializing the network with parameters from a pretrained GoogLeNet [<xref ref-type="bibr" rid="ref-29">29</xref>] and, combined with data augmentation, achieved an accuracy of 86.1% on a Twitter image dataset.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Multimodal Sentiment Analysis</title>
<p>Although previous studies have demonstrated the effectiveness of unimodal approaches for emotion recognition, comparative analyses in the literature consistently show that multimodal strategies achieve superior performance [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
<p>MSA seeks to infer sentiment by integrating heterogeneous modalities&#x2014;such as vision, text, and audio&#x2014;to provide a more comprehensive understanding of human affect. The recent surge in multimodal content has driven renewed interest in MSA. You et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] proposed a Cross-modal Consistency Regression (CCR) model that uses CNNs to extract visual features and a paragraph-vector model to obtain textual features; consistency across modalities is then enforced via Kullback&#x2013;Leibler (KL) divergence to enable joint optimization. Zhao et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] introduced an image&#x2013;text consistency&#x2013;driven approach that first uses an SVM classifier to assess whether image and text semantics align, and then combines intermediate visual features from SentiBank [<xref ref-type="bibr" rid="ref-25">25</xref>] with textual and social features for adaptive MSA. To strengthen inter-modal interaction, several works employ attention mechanisms. Zhang et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed an attention-based model with a symmetric architecture that processes text and image data in parallel: a denoising autoencoder extracts robust textual features, while an improved variational autoencoder with attention extracts salient visual features. They further design a cross-feature fusion module that uses attention to learn complementary information between text and image features rather than relying on simple concatenation or one-way influence. Zhou et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] introduced CAHFW-Net, a method for video-based emotion recognition that combines cross-attention with a hybrid feature-weighted neural network. Its hierarchical attention encoding network&#x2014;particularly the cross-attention module&#x2014;effectively captures complementary cues between facial expressions and scene context, mitigating emotion confusion and misinterpretation. Khan and Fu [<xref ref-type="bibr" rid="ref-35">35</xref>] developed a dual-stream model based on input-space translation (EF-CapTrBERT): an object-aware Transformer first translates images into natural-language descriptions, which are then formed into auxiliary sentences to inject multimodal information into BERT. This approach addresses image noise and modality fusion without modifying BERT&#x2019;s internal architecture. With the rapid progress of vision&#x2013;language models (VLMs) [<xref ref-type="bibr" rid="ref-36">36</xref>], their strong multimodal capabilities have seen wide adoption. For example, Huang et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] proposed CLIP-MSA, a CLIP-based MSA model that introduces CLIP&#x2019;s cross-modal encoder as the second branch of a dual-branch feature extractor. By leveraging CLIP&#x2019;s pretrained semantic alignment and external knowledge, CLIP-MSA substantially improves the quality of unimodal representations.</p>
<p>In summary, although existing methods have achieved notable progress on MSA, they still exhibit clear limitations in hierarchical modeling of visual&#x2013;textual features and in the fine-grained characterization of dynamic inter-modal interactions. Most current approaches follow a single-path feature extraction paradigm and perform multimodal fusion via simple concatenation or summation. Such coarse-grained modeling is insufficient to fully explore discriminative intra-modal representations from multiple perspectives, and it fails to effectively capture the complementarity and alignment between images and text in a fine-grained semantic space, which often leads to suboptimal sentiment recognition in complex real-world scenarios. Recent CLIP-based methods (e.g., FAMEAC [<xref ref-type="bibr" rid="ref-10">10</xref>], SentiCLIP [<xref ref-type="bibr" rid="ref-38">38</xref>]) and large vision-language models have made important advances; however, they have not yet simultaneously achieved multi-scale hierarchical representation, bidirectional fine-grained cross-modal alignment, and adaptive dynamic fusion in a unified framework. To bridge these remaining gaps, the proposed MB-CMCAG network performs multi-path collaborative modeling to learn hierarchical visual&#x2013;textual representations at both coarse and fine granularities, employs bidirectional cross-modal cross-attention for deep semantic interaction, and introduces a novel cross-modal interactive gating mechanism that adaptively selects and reweights cross-modal information, thereby effectively suppressing redundant noise and highlighting key cues critical for sentiment discrimination.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<sec id="s3_1">
<label>3.1</label>
<title>Task Definition</title>
<p>Given a multimodal sentiment analysis dataset <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, each sample consists of an image <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, an original text <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, and a sentiment label <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mtext>positive</mml:mtext><mml:mo>,</mml:mo><mml:mtext>neutral</mml:mtext><mml:mo>,</mml:mo><mml:mtext>negative</mml:mtext><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. To better exploit the semantic and emotional cues embedded in the visual modality, an image caption generation model <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is employed to generate a descriptive caption <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for each image. The generated caption is concatenated with the original text to form the final textual input <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The goal of this task is to learn a multimodal sentiment classification mapping function <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>f</mml:mi><mml:mo>:</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mi>y</mml:mi></mml:math></inline-formula> based on the image <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and the augmented text <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. By incorporating image-generated captions, the model can jointly learn complementary emotional representations from both visual and linguistic modalities, thereby achieving more accurate multimodal sentiment understanding.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Model Overview</title>
<p>The proposed MB-CMCAG model consists of five main components: an image caption generation module, a multi-branch feature-extraction layer, a cross-modal cross-attention interaction layer, a cross-modal cross-gating fusion layer, and a sentiment classification layer. As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the image captioning module converts an image into a textual description, which is then concatenated with the original text to form the final composite text. The multi-branch feature-extraction module leverages different pretrained models to obtain multi-granularity feature representations from both images and text. The cross-modal cross-attention interaction module enhances information exchange between modalities through bidirectional attention, i.e., image-to-text and text-to-image. The cross-modal cross-gating fusion module adaptively controls the fusion strength of features from different modalities and integrates contextual information. Finally, the fused multimodal representation is fed into a multi-layer classifier for sentiment prediction.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overall architecture of the MB-CMCAG model.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81626-fig-1.tif"/>
</fig>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Multi-Branch Feature Extraction</title>
<p>In multimodal sentiment analysis, our inputs consist of two primary modalities: an image <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and a text <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is an augmented text formed by concatenating the original text <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> with the image-generated caption <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. To comprehensively capture the characteristics of both modalities, we design a multi-branch feature extraction module.</p>
<p>Unlike single-branch encoding strategies that rely on a unified representation space, the proposed dual-branch design aims to explicitly model complementary semantic information at different granularity levels, thereby mitigating the limitations of either overly coarse or overly local feature representations.</p>
<p>Specifically, the fine-grained branch employs ViT to extract patch-level local features from the image and BERT to perform token-level semantic analysis on the enhanced text, thereby preserving rich spatial structures and sequential details for fine-grained semantic alignment between modalities. The coarse-grained branch leverages a pre-trained CLIP model to extract global semantic embeddings for both image and text, providing high-level cross-modal consistency alignment. However, CLIP&#x2019;s highly compressed global vectors inevitably discard the original spatial layout and sequential structure. Therefore, the proposed multi-branch architecture is not a simple stacking of pre-trained models; instead, it deliberately constructs two complementary representation perspectives&#x2014;the fine-grained branch captures local discriminative information while the coarse-grained branch supplies robust global semantic context&#x2014;forming a richer, more hierarchical feature space that provides a sufficient and complementary foundation for the subsequent cross-modal cross-attention module and dynamic gating mechanism.</p>
<p>Visual Feature Extraction Using ViT</p>
<p>For the input image <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, we employ a ViT to extract visual features. The ViT divides the image into fixed-size patches, linearly embeds each patch, and adds positional encodings to form a sequence that is fed into the Transformer encoder. Concretely, the image <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is split into <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>P</mml:mi></mml:math></inline-formula> equal-sized patches; each patch is mapped via a linear projection into a <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula>-dimensional feature space, producing the image feature sequence <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ViT</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> denotes the feature vector of the <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>j</mml:mi></mml:math></inline-formula>-th image patch, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>P</mml:mi></mml:math></inline-formula> is the total number of patches, and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> is the feature dimensionality of the ViT output.</p>
<p>Text Feature Extraction Using BERT</p>
<p>For the textual input <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the text is first tokenized and then encoded using the BERT model. Owing to its bidirectional Transformer architecture, BERT effectively captures contextual information in the text, thereby producing richer textual representations. The text sequence <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is encoded by BERT as follows:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>BERT</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> denotes the contextual semantic feature of the <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>j</mml:mi></mml:math></inline-formula>-th token, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>L</mml:mi></mml:math></inline-formula> is the length of the text sequence, and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula> is the feature dimensionality of the BERT output.</p>
<p>Multimodal Feature Extraction Using CLIP</p>
<p>In order to enhance the model&#x2019;s understanding of cross-modal semantic relations, we incorporate the Contrastive Language&#x2013;Image Pretraining (CLIP) model as an additional pathway to obtain joint image&#x2013;text representations. CLIP uses contrastive learning to project images and text into a shared embedding space, where corresponding image&#x2013;text pairs are pulled closer, thereby providing globally aligned semantic representations. Images and text are processed by CLIP&#x2019;s image encoder and text encoder, respectively. For a given image <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and text <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, we compute the outer product (<inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula>) between their CLIP features to capture cross-modal interactions beyond simple concatenation, followed by a diagonalization operation to reduce redundancy and retain the most informative aligned components:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>image</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msubsup><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>text</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Diag</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:msubsup><mml:mo>&#x2297;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> denotes the output dimensionality of the CLIP model, <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mo>&#x2297;</mml:mo></mml:math></inline-formula> denotes the outer product, and <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mi mathvariant="normal">D</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">g</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is applied to extract the diagonal values.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Multimodal Feature Fusion Module</title>
<p>In the image&#x2013;text feature fusion module, feature fusion remains one of the core bottlenecks that constrain further performance improvement. Because the visual modality (e.g., facial expressions and scene composition) and the textual modality (e.g., dialogue semantics and affective words) exhibit substantial heterogeneity in representation form, information density, and spatiotemporal characteristics, conventional fusion strategies&#x2014;such as concatenation, weighted averaging, or simple attention mechanisms&#x2014;are insufficient to achieve deep semantic interaction. To address this limitation, we propose a bidirectional cross-modal attention mechanism to overcome the technical bottleneck, as illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. By constructing dual information flow paths from image to text and from text to image, the model enables deep and collaborative modeling of cross-modal features.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The cross-modal attention mechanism.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81626-fig-2.tif"/>
</fig>
<p>Image-to-Text Cross-Modal Attention</p>
<p>In the Image-to-Text cross-modal attention module, image features are used as queries, while textual features serve as keys and values. This design enables the model to attend to text regions that are relevant to the visual content. The computation of the image-to-text cross-modal multi-head attention follows the scaled dot-product attention formulation proposed by Vaswani et al. [<xref ref-type="bibr" rid="ref-39">39</xref>] and is formulated as follows:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>Q</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>V</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>MultiHead</mml:mtext></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Concat</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>O</mml:mi></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>Q</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>V</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are learnable projection matrices, and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mn>2</mml:mn><mml:mi>T</mml:mi></mml:mrow><mml:mi>O</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mtext>heads</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the output projection matrix. Here, <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:math></inline-formula> denotes the dimensionality of the attention mechanism.</p>
<p>Text-to-Image Cross-Modal Attention</p>
<p>In the Text-to-Image cross-modal attention module, textual features are used as queries, while visual features serve as keys and values. This design enables the model to attend to image regions that are relevant to the textual content. The computation of the text-to-image multi-head attention mechanism is formulated as follows:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>Q</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi>T</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>V</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>MultiHead</mml:mtext></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Concat</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>O</mml:mi></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>Q</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>K</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>V</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> are learnable projection matrices, and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>T</mml:mi><mml:mn>2</mml:mn><mml:mi>I</mml:mi></mml:mrow><mml:mi>O</mml:mi></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mtext>heads</mml:mtext></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> is the output projection matrix.</p>
<p>Cross-Attention over CLIP Features</p>
<p>Building upon CLIP-based feature extraction, this module introduces a cross-attention mechanism to enable deep semantic interaction between the visual and textual modalities. Previously, we captured second-order interactions between image and text features via the tensor outer product and extracted key auto-correlation terms through a diagonalization operation, thereby constructing a compact bilinear representation. On this basis, cross-attention further aggregates and recalibrates cross-modal information, producing a fused representation with higher information density and stronger discriminative power.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>MultiHeadAttn</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>W</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>b</mml:mi><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the learnable projection matrix and bias term, respectively, and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>d</mml:mi></mml:math></inline-formula> is the hidden dimension of the attention module.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Cross-Modal Interactive Gating Mechanism</title>
<p>During multimodal feature fusion&#x2014;particularly between image and text modalities&#x2014;the contribution of each modality to final sentiment polarity typically exhibits marked imbalance. This imbalance is compounded by inter-modal redundancy, noise, and cross-modal feature misalignment, which undermine conventional fusion strategies and can lead to loss of complementary information or suboptimal representations. To address this issue, we propose a cross-modal gated fusion mechanism that dynamically adjusts the contribution of each modality according to the current context, as illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The mechanism introduces learnable attention-gating units that adaptively adjust fusion weights during training, selectively amplifying sentiment-relevant shared signals while suppressing redundant or noisy components.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The cross-modal gated mechanism.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81626-fig-3.tif"/>
</fig>
<p>Concretely, the gating mechanism first computes a gating signal from the image&#x2013;text features that reflects the relative importance of each modality in the current context. The computed gate is then used to weight the features, and the gated outputs are applied via element-wise modulation to the target modality feature maps. Finally, the gated features are concatenated and integrated to form a more discriminative multimodal joint representation. This adaptive gating strategy increases fusion flexibility and partially mitigates modality misalignment, yielding more robust and information-rich features for subsequent sentiment classification.</p>
<p>To adaptively fuse the global semantic representation with bidirectional cross-modal interaction features, we design a gated fusion module. Given <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and a cross-modal feature <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, the gated fusion is defined as follows.</p>
<p>First, three transformation branches are constructed:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mi>X</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mi>X</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mi>X</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>W</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>b</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters.</p>
<p>Then, the joint gated interaction feature is obtained via element-wise multiplication:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>Z</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>X</mml:mi></mml:msub><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> denotes the Hadamard product.</p>
<p>To further enhance representation capability and stabilize optimization, a residual refinement mapping is applied:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mi>Z</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub><mml:mi>Z</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>W</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters.</p>
<p>Accordingly, the gated fusion outputs for the two cross-modal directions are computed as:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mrow><mml:mi>&#x1D4A2;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">C</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">I</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the gated fusion function defined by <xref ref-type="disp-formula" rid="eqn-12">Eqs. (12)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-16">(16)</xref>.</p>
<p>Finally, the two gated features are concatenated and projected to form the final fused representation:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mover><mml:mi>F</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mi>W</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>H</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> denotes the concatenation operation, <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>W</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>b</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are learnable parameters.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Sentiment Classification and Objective</title>
<p>Finally, given the fused multimodal representation <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mrow><mml:mtext mathvariant="bold">F</mml:mtext></mml:mrow></mml:math></inline-formula>, we feed it into a classifier to predict the sentiment label <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>. The model adopts a fully connected layer followed by a Softmax function for multi-class classification, and is trained with the cross-entropy loss <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The entire network is optimized in an end-to-end manner by minimizing this loss.
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mtext mathvariant="bold">F</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>CrossEntropyLoss</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mrow><mml:mtext mathvariant="bold">W</mml:mtext></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>b</mml:mi></mml:math></inline-formula> are learnable parameters.</p>
</sec>
<sec id="s3_2_5">
<label>3.2.5</label>
<title>Image Caption Generation Module</title>
<p>To more comprehensively extract visual information, we employ an image captioning module adapted from the Caption Transformer (CaTr) architecture proposed by Khan and Fu [<xref ref-type="bibr" rid="ref-35">35</xref>]. This module converts the raw image into a natural language description, enabling seamless integration with text-based language models without modifying the core architecture.</p>
<p>Specifically, the CaTr module employs ResNet-101 [<xref ref-type="bibr" rid="ref-40">40</xref>] as the convolutional backbone to extract visual features from the input image <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>I</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. These features are subsequently projected to a lower-dimensional embedding space of <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>256</mml:mn></mml:math></inline-formula> via a <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolutional layer. Fixed positional encodings are then added to the resulting feature map, which is passed through a stack of Transformer encoder layers adapted from the DETR architecture [<xref ref-type="bibr" rid="ref-41">41</xref>]. The encoder leverages multi-head self-attention to capture object-level dependencies across the image. Finally, the decoder produces a natural language description of the image through non-autoregressive generation.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>To evaluate the performance of the proposed MB-CMCAG model, this study utilizes two widely adopted Multi-View Sentiment Analysis (MVSA) datasets: MVSA-Single and MVSA-Multiple [<xref ref-type="bibr" rid="ref-42">42</xref>]. Both datasets are sourced from Twitter and consist of image-text pairs, each annotated with sentiment labels: positive, neutral, and negative. The MVSA-Single dataset contains 5129 image-text pairs, each labeled by a single annotator. After preprocessing to remove samples with inconsistent sentiment labels across modalities, 4511 valid samples remain, including 2683 positive samples, 470 neutral samples, and 1358 negative samples. The MVSA-Multiple dataset comprises 19,600 image-text pairs, labeled by three annotators. The final sentiment label is determined through majority voting. After preprocessing, 17,024 valid samples are obtained, including 11,318 positive samples, 4408 neutral samples, and 1298 negative samples. To facilitate comparisons with other models, both datasets are randomly split into training, validation, and test sets in an 8:1:1 ratio. <xref ref-type="table" rid="table-1">Table 1</xref> provides detailed information on the dataset distribution.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Detailed statistics of the MVSA-single and MVSA-multiple datasets.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Split</th>
<th>Positive</th>
<th>Neutral</th>
<th>Negative</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>MVSA-Single</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td></td>
<td>Train</td>
<td>2146</td>
<td>376</td>
<td>1086</td>
<td>3608</td>
</tr>
<tr>
<td></td>
<td>Valid</td>
<td>268</td>
<td>47</td>
<td>135</td>
<td>450</td>
</tr>
<tr>
<td></td>
<td>Test</td>
<td>269</td>
<td>47</td>
<td>137</td>
<td>453</td>
</tr>
<tr>
<td></td>
<td>Total</td>
<td>2683</td>
<td>470</td>
<td>1358</td>
<td>4511</td>
</tr>
<tr>
<td><bold>MVSA-Multiple</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td></td>
<td>Train</td>
<td>9054</td>
<td>3526</td>
<td>1038</td>
<td>13,618</td>
</tr>
<tr>
<td></td>
<td>Valid</td>
<td>1131</td>
<td>440</td>
<td>129</td>
<td>1700</td>
</tr>
<tr>
<td></td>
<td>Test</td>
<td>1133</td>
<td>442</td>
<td>131</td>
<td>1706</td>
</tr>
<tr>
<td></td>
<td>Total</td>
<td>11,318</td>
<td>4408</td>
<td>1298</td>
<td>17,024</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Settings and Hyperparameters</title>
<p>All experiments were conducted on a server equipped with an NVIDIA RTX 4090 GPU (24 GB VRAM). The model was implemented in PyTorch. All experimental results are averaged over three runs with different random seeds. For feature extraction of the original text and image captions, we used a pretrained BERT tokenizer with a maximum sequence length of 128; sequences longer than this were truncated and shorter ones were padded. For visual feature extraction, we employed an ImageNet-pretrained ViT-B/16 model with a patch size of <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mn>16</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula>. For the additional branch, we used the CLIP-ViT-B/32 pretrained model to extract complementary semantic features. Models were trained with the Adam optimizer using a learning rate of <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and a weight decay of <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. Owing to differences in dataset sizes, the batch size was set to 32 for MVSA-Single and 128 for MVSA-Multiple. Training proceeded for up to 30 epochs with early stopping based on validation performance and a patience of 5 epochs. Detailed hyperparameter settings are reported in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Hyperparameter settings.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Learning_Rate</td>
<td>5e&#x2212;5</td>
</tr>
<tr>
<td>Optimizer</td>
<td>Adam</td>
</tr>
<tr>
<td>Batch size</td>
<td>32/128</td>
</tr>
<tr>
<td>Epochs</td>
<td>30</td>
</tr>
<tr>
<td>Dropout</td>
<td>0.3</td>
</tr>
<tr>
<td>Weight decay</td>
<td>1e&#x2212;4</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In terms of evaluation metrics, we employ Accuracy and Macro-F1 as the primary measures, which are standard metrics for multi-class sentiment analysis tasks. Accuracy measures the overall proportion of correctly classified samples, while Macro-F1 provides a balanced evaluation of class-wise precision and recall, making it more suitable for datasets with class imbalance. The calculation formulas are as follows:<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula>
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula>
<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:mi>F</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:msub><mml:mi>l</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></disp-formula>
<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mi>F</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>c</mml:mi></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>N</mml:mi></mml:math></inline-formula> denotes the total number of test samples, <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> denote the predicted label and the ground-truth label of the <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>i</mml:mi></mml:math></inline-formula>-th sample, respectively, <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the indicator function, and <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow></mml:math></inline-formula> denotes the set of sentiment classes. For each class <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>c</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>F</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>F</mml:mi><mml:msub><mml:mi>N</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:math></inline-formula> are computed in a one-vs.-rest manner.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Model Comparisons and Baselines</title>
<p>To demonstrate the effectiveness of the proposed MB-CMCAG model, we conduct comparative experiments against representative baseline methods from the literature. The performance differences are quantitatively evaluated from both unimodal and multimodal perspectives using two core metrics, namely Accuracy and Macro-F1. The comparative results are reported in <xref ref-type="table" rid="table-3">Table 3</xref>. The baseline methods included in the comparison are described as follows:</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison with other models.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Modality</th>
<th align="center" rowspan="2">Method</th>
<th colspan="2">MVSA-Single</th>
<th colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th>ACC</th>
<th>F1</th>
<th>ACC</th>
<th>F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Text</td>
<td>BERT</td>
<td>71.31</td>
<td>69.71</td>
<td>67.57</td>
<td>66.27</td>
</tr>
<tr>
<td></td>
<td>BiLSTM</td>
<td>70.14</td>
<td>65.11</td>
<td>67.52</td>
<td>66.82</td>
</tr>
<tr>
<td>Image</td>
<td>ViT</td>
<td>63.95</td>
<td>62.44</td>
<td>62.12</td>
<td>61.24</td>
</tr>
<tr>
<td></td>
<td>ResNet-50</td>
<td>64.72</td>
<td>61.64</td>
<td>61.93</td>
<td>61.05</td>
</tr>
<tr>
<td>Text and Image</td>
<td>MultiSentiNet</td>
<td>69.84</td>
<td>69.63</td>
<td>68.86</td>
<td>68.11</td>
</tr>
<tr>
<td></td>
<td>Co-Memory</td>
<td>70.16</td>
<td>70.43</td>
<td>69.78</td>
<td>69.94</td>
</tr>
<tr>
<td></td>
<td>DMAF</td>
<td>71.87</td>
<td>71.59</td>
<td>70.47</td>
<td>70.38</td>
</tr>
<tr>
<td></td>
<td>MVAN</td>
<td>73.08</td>
<td>73.12</td>
<td>72.42</td>
<td>72.35</td>
</tr>
<tr>
<td></td>
<td>ITIN</td>
<td>75.14</td>
<td>74.77</td>
<td>73.38</td>
<td><bold>73.26</bold></td>
</tr>
<tr>
<td></td>
<td>CLMLF</td>
<td>75.33</td>
<td>73.45</td>
<td>72.00</td>
<td>69.83</td>
</tr>
<tr>
<td></td>
<td>CLIP-CA-CG</td>
<td>75.36</td>
<td>75.18</td>
<td>73.56</td>
<td>73.79</td>
</tr>
<tr>
<td></td>
<td>SentiCLIP</td>
<td>76.05</td>
<td>75.27</td>
<td>73.32</td>
<td>70.45</td>
</tr>
<tr>
<td></td>
<td>VILA</td>
<td>76.21</td>
<td>75.39</td>
<td>73.26</td>
<td>72.91</td>
</tr>
<tr>
<td></td>
<td>LLaVAC</td>
<td>75.82</td>
<td>74.91</td>
<td>73.15</td>
<td>72.68</td>
</tr>
<tr>
<td></td>
<td><bold>MB-CMCAG</bold></td>
<td><bold>76.38</bold></td>
<td><bold>75.46</bold></td>
<td><bold>73.87</bold></td>
<td>72.64</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-3fn1" fn-type="other">
<p>Note: &#x002A;indicates statistically significant improvement over the strongest baseline in the corresponding column based on three independent runs (paired <italic>t</italic>-test, <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>p</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mn>0.05</mml:mn></mml:math></inline-formula>). Bold denotes the best result. All results are averaged over three independent runs with different random seeds.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><bold>BERT and BiLSTM [<xref ref-type="bibr" rid="ref-22">22</xref>]:</bold> BERT and BiLSTM represent classical methods for modeling textual modalities, with extensive applications in text classification tasks.</p>
<p><bold>ViT and ResNet-50 [<xref ref-type="bibr" rid="ref-40">40</xref>]:</bold> ResNet-50 and ViT represent the most prevalent benchmark models in image classification.</p>
<p><bold>MultiSentiNet [<xref ref-type="bibr" rid="ref-43">43</xref>]:</bold> This model introduces visual semantic features (objects and scenes) as auxiliary information and employs a visually guided attention LSTM to mitigate semantic loss and weak image-text correlation caused by naive feature fusion.</p>
<p><bold>Co-Memory [<xref ref-type="bibr" rid="ref-44">44</xref>]:</bold> This work addresses insufficient image-text interaction in multimodal sentiment analysis by proposing a co-memory network, where text guides visual feature selection and images attend to key words, with stacked layers for iterative refinement.</p>
<p><bold>DMAF [<xref ref-type="bibr" rid="ref-45">45</xref>]:</bold> This model uses unimodal attention to highlight sentiment regions and words, and integrates intermediate and late fusion for robust multimodal sentiment prediction.</p>
<p><bold>MVAN [<xref ref-type="bibr" rid="ref-46">46</xref>]:</bold> This paper proposes a Multi-View Attention Network that jointly models object-level and scene-level views of images and employs an interactive learning mechanism to improve multimodal sentiment analysis performance.</p>
<p><bold>CLMLF [<xref ref-type="bibr" rid="ref-47">47</xref>]:</bold> This method employs a Transformer-based multi-layer fusion module to align and fuse textual and visual features at the token level, and incorporates label- and data-driven contrastive learning to enhance the learning of sentiment-discriminative shared representations.</p>
<p><bold>CLIP-CA-CG [<xref ref-type="bibr" rid="ref-48">48</xref>]:</bold> This author builds a CLIP-based cross-modal sentiment model that extracts visual and textual features with ResNet50 and RoBERTa, enhances cross-modal interaction via multi-head attention, and fuses multi-level features using a cross-modal gating module.</p>
<p><bold>SentiCLIP [<xref ref-type="bibr" rid="ref-38">38</xref>]:</bold> This multimodal sentiment analysis model leverages the CLIP encoder to construct a unified semantic space and employs a feature interaction module to enhance cross-modal emotion understanding.</p>
<p><bold>ITIN [<xref ref-type="bibr" rid="ref-49">49</xref>]:</bold> The paper proposes an image-text interaction network that enhances multimodal sentiment analysis accuracy through fine-grained alignment of image regions and text words, along with adaptive feature fusion.</p>
<p><bold>VILA [<xref ref-type="bibr" rid="ref-50">50</xref>]:</bold> The paper significantly improves visual language model performance and surpasses LLaVA-1.5 by unfreezing the LLM, adopting interleaved data, and using mixed text instruction fine-tuning.</p>
<p><bold>LLaVAC [<xref ref-type="bibr" rid="ref-51">51</xref>]:</bold> The paper designs a structured prompt incorporating independent image and text sentiment labels as well as a joint multimodal label, and performs single-round LoRA fine-tuning on LLaVA to directly use it as a multimodal sentiment classifier.</p>
<p>The results in <xref ref-type="table" rid="table-3">Table 3</xref> show that MB-CMCAG achieves statistically significant improvements over the strongest baselines on the corresponding metrics, further confirming the robustness of the proposed method. Although MB-CMCAG achieves the best accuracy on MVSA-Multiple, its F1-score is slightly lower than that of some comparison methods. This is mainly because MVSA-Multiple is a more challenging dataset, where image-text semantic inconsistencies, ambiguous sentiment expressions, and class imbalance are more pronounced. Since F1-score is more sensitive to minority-class recognition than Accuracy, a small fluctuation in difficult categories may lead to a lower F1 even when overall accuracy remains strong. Therefore, this result does not contradict the effectiveness of MB-CMCAG, but rather reflects the difficulty of the dataset and the remaining room for improvement in minority-class discrimination.</p>

<p>Furthermore, to more intuitively illustrate the discriminative capability of the MB-CMCAG model across different categories, we plot the corresponding confusion matrices on the MVSA-Single and MVSA-Multiple datasets, respectively. The confusion matrices indicate that the model achieves higher classification accuracy on positive and negative samples, whereas its performance on the neutral category is relatively weaker, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. This observation is consistent with the slightly lower F1-score on MVSA-Multiple, since errors in minority classes have a larger impact on Macro-F1 than on overall Accuracy. Future work may explore targeted data augmentation and cost-sensitive learning to further improve neutral-class recognition and alleviate class imbalance.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Confusion matrices of the MB-CMCAG model on the MVSA-Single and MVSA-Multiple datasets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81626-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Ablation Study</title>
<p>To systematically evaluate the contribution of each core component in the proposed MB-CMCAG, we conducted a seven-group ablation study on both the MVSA-Single and MVSA-Multiple datasets. In each variant, one key module is removed or replaced while all other components remain unchanged. The seven settings are: (1) Text Only, (2) Image Only, (3) MB-CMCAG without CLIP (w/o CLIP), (4) MB-CMCAG without BERT-ViT (w/o BV), (5) MB-CMCAG without Cross-Modal Cross-Attention (w/o CMCA), (6) MB-CMCAG without Gated Cross-Attention (w/o GCA), and (7) MB-CMCAG without Caption Generation Module (w/o CGM). Performance is evaluated using Accuracy and Macro-F1. The detailed experimental settings and complete results are reported in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation studies of the MB-CMCAG model.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Model</th>
<th align="center" colspan="2">MVSA-Single</th>
<th align="center" colspan="2">MVSA-Multiple</th>
</tr>
<tr>
<th>Accuracy</th>
<th>F1</th>
<th>Accuracy</th>
<th>F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Text Only</td>
<td>71.68</td>
<td>70.56</td>
<td>69.12</td>
<td>68.56</td>
</tr>
<tr>
<td>Image Only</td>
<td>69.97</td>
<td>69.05</td>
<td>68.64</td>
<td>68.17</td>
</tr>
<tr>
<td>MB-CMCAG w/o CLIP</td>
<td>73.85</td>
<td>72.46</td>
<td>71.22</td>
<td>69.59</td>
</tr>
<tr>
<td>MB-CMCAG w/o BV</td>
<td>72.77</td>
<td>71.98</td>
<td>70.29</td>
<td>69.32</td>
</tr>
<tr>
<td>MB-CMCAG w/o CMCA</td>
<td>72.48</td>
<td>70.53</td>
<td>70.74</td>
<td>69.18</td>
</tr>
<tr>
<td>MB-CMCAG w/o GCA</td>
<td>74.81</td>
<td>72.03</td>
<td>72.58</td>
<td>70.39</td>
</tr>
<tr>
<td>MB-CMCAG w/o CGM</td>
<td>75.79</td>
<td>74.82</td>
<td>73.06</td>
<td>71.96</td>
</tr>
<tr>
<td>MB-CMCAG</td>
<td>76.38</td>
<td>75.46</td>
<td>73.87</td>
<td>72.64</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The ablation results clearly show that every module contributes positively to the final performance. Removing any single component leads to a consistent decrease in both Accuracy and Macro-F1 on the two datasets, confirming the necessity and complementarity of the proposed design. Specifically, the unimodal variants (Text Only and Image Only) suffer the largest performance drops, with average accuracy decreases of 4.73% and 5.82%, respectively, highlighting the importance of multimodal fusion. Removing the CLIP branch (w/o CLIP) or the fine-grained BV branch (w/o BV) reduces average accuracy by 2.59% and 3.59%, respectively, indicating that coarse-grained global alignment and fine-grained local modeling are both essential. Eliminating the cross-modal cross-attention module (w/o CMCA) causes a 3.52% drop, showing that bidirectional fine-grained interaction is critical for semantic alignment. Replacing the gated mechanism with simple concatenation (w/o GCA) leads to a 1.43% decline, demonstrating the effectiveness of adaptive noise suppression and feature re-weighting. Finally, removing the image caption generation module (w/o CGM) results in a 0.70% decrease, confirming that visual-to-textual semantic enrichment provides useful complementary information. Overall, these results verify that the performance gains of MB-CMCAG arise from the coordinated interaction of its modules rather than from any single component alone, which further supports the effectiveness of the proposed hierarchical design.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Computational Efficiency Analysis</title>
<p>To further evaluate the computational efficiency of the proposed method, we report the model size, computational cost in terms of Floating Point Operations per Second (FLOPs), inference time, and GPU memory usage in <xref ref-type="table" rid="table-5">Table 5</xref>. All measurements are conducted on the same server under the same settings as <xref ref-type="sec" rid="s4_2">Section 4.2</xref>, with a batch size of 1 to ensure fair inference time comparison.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Computational complexity and efficiency analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model Variant</th>
<th>Parameters (M)</th>
<th>FLOPs (G)</th>
<th>Inference Time (ms)</th>
<th>GPU Memory (GB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>BERT only</td>
<td>110</td>
<td>11.2</td>
<td>18.4</td>
<td>2.1</td>
</tr>
<tr>
<td>ViT-B/16 only</td>
<td>86</td>
<td>15.8</td>
<td>22.6</td>
<td>3.4</td>
</tr>
<tr>
<td>CLIP-ViT-B/32 only</td>
<td>151</td>
<td>12.5</td>
<td>19.7</td>
<td>2.8</td>
</tr>
<tr>
<td>MB-CMCAG (full model)</td>
<td>347</td>
<td>39.6</td>
<td>45.3</td>
<td>6.9</td>
</tr>
<tr>
<td>LLaVAC</td>
<td>7000</td>
<td>520</td>
<td>248</td>
<td>14.2</td>
</tr>
<tr>
<td>VILA</td>
<td>7500</td>
<td>580</td>
<td>265</td>
<td>15.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Although MB-CMCAG integrates three pre-trained backbones, its total parameter count (347M) and inference time (45.3 ms) remain significantly lower than recent large vision-language models such as LLaVAC and VILA (both over 7B parameters). The cross-modal gated mechanism effectively suppresses redundant computation, resulting in a practical trade-off between performance and efficiency. Compared with these large-scale VLMs, our model achieves competitive accuracy while requiring only about 5% of the parameters and less than 20% of the inference time, demonstrating its suitability for real-world deployment on standard GPUs.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Case Study</title>
<p>To further illustrate the effectiveness of MB-CMCAG in multimodal text-image classification tasks, we selected four representative cases from the MVSA-Single test set. These examples are used to qualitatively analyze the model&#x2019;s ability to capture subtle emotional cues in text-image pairs. As shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, we compare our proposed model with CLIP-CA-CG and SentiCLIP on these four cases. In addition, we input image captions as auxiliary sentences to further examine their effect on sentiment prediction.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Classification case studies of CLIP-CA-CG, SentiCLIP, and MB-CMCAG on the MVSA dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81626-fig-5.tif"/>
</fig>
<p>First, all three models correctly classify the first and second examples as positive and negative, respectively, since the image-text pairs in these cases convey clear and consistent emotional signals. In the third example, where both modalities express no obvious sentiment tendency, CLIP-CA-CG predicts negative and SentiCLIP predicts positive, whereas our model correctly identifies the sample as neutral. This result suggests that the proposed multi-branch feature extraction strategy can better preserve complementary semantic information from both modalities, thereby improving the recognition of neutral samples with ambiguous emotional cues. In the fourth example, both CLIP-CA-CG and SentiCLIP classify the sample as positive, mainly because the image itself does not explicitly express a negative tendency while the text suggests positive sentiment. In contrast, our model correctly predicts negative, which indicates that the image captioning module can provide additional visual semantic cues and help the model handle image-text inconsistency more effectively.</p>
<p>In summary, these representative cases show that MB-CMCAG can better capture complementary sentiment cues from images and texts, especially in challenging samples with weak or inconsistent modality alignment. The results also indicate that image captions as auxiliary textual descriptions can further enhance cross-modal understanding and improve sentiment classification in complex scenarios.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this study, we propose a multimodal sentiment analysis framework, MB-CMCAG, based on a multi-branch cross-modal cross-attention gated network. The framework addresses several key challenges in feature extraction, cross-modal alignment, and fusion. It first converts visual information into supplementary text via an image captioning model, thereby enriching the semantic representation of the original textual input. Next, it employs a dual-path parallel feature extraction architecture: one path leverages ViT and BERT for fine-grained visual and textual modeling, while the other exploits a pre-trained CLIP model to extract cross-modal coarse-grained features and adopts its diagonal values as an efficient semantic representation. Cross-modal cross-attention and multi-head attention modules are then specifically designed for each path to strengthen inter-modal semantic alignment. Finally, a cross-modal interaction gated fusion mechanism integrates the features from both paths, which are subsequently fed into a classification layer for sentiment prediction.</p>
<p>Extensive experiments on the MVSA benchmark datasets show that MB-CMCAG achieves competitive performance compared to existing methods in terms of accuracy and F1 score. Although the framework demonstrates promising results, future work may extend it to additional modalities such as video and audio, enabling the capture of richer multidimensional emotional expressions.</p>
</sec>
</body>
<back>
<ack>
<p>This work was supported by the National Natural Science Foundation of China and the Open Fund of the Guangxi Key Laboratory of Digital Infrastructure.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research project was supported by the National Natural Science Foundation of China (62266004) and the Open Fund of the Key Laboratory of Digital Infrastructure in Guangxi (GXDINBC202401).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Xinshan Huang conceived the idea, designed the proposed model, implemented the framework, conducted the experiments, and drafted the manuscript. Zirui Pei and Chaohong Tan assisted with experimental validation and manuscript revision. Zuqiang Meng provided conceptual guidance and funding support. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The MVSA dataset used in this study is publicly available and can be accessed through [<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Leng</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Deep learning-based multimodal emotion recognition from audio, visual, and text modalities: A systematic review of recent advancements and future prospects</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>237</volume>:<fpage>121692</fpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chowdary</surname> <given-names>MK</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>TN</given-names></string-name>, <string-name><surname>Hemanth</surname> <given-names>DJ</given-names></string-name></person-group>. <article-title>Deep learning-based facial emotion recognition for human-computer interaction applications</article-title>. <source>Neural Comput Appl</source>. <year>2023</year>;<volume>35</volume>(<issue>32</issue>):<fpage>23311</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-021-06012-8</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Multimodal sentiment analysis based on fusion methods: A survey</article-title>. <source>Inform Fus</source>. <year>2023</year>;<volume>95</volume>:<fpage>306</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2023.02.028</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pandey</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vishwakarma</surname> <given-names>DK</given-names></string-name></person-group>. <article-title>Progress, achievements, and challenges in multimodal sentiment analysis using deep learning: A survey</article-title>. <source>Appl Soft Comput</source>. <year>2024</year>;<volume>152</volume>:<fpage>111206</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2023.111206</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Multimodal emotion recognition with temporal and semantic consistency</article-title>. <source>IEEE ACM Trans Audio Speech Lang Process</source>. <year>2021</year>;<volume>29</volume>:<fpage>3592</fpage>&#x2013;<lpage>603</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taslp.2021.3129331</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dixit</surname> <given-names>C</given-names></string-name>, <string-name><surname>Satapathy</surname> <given-names>SM</given-names></string-name></person-group>. <article-title>Deep CNN with late fusion for real time multimodal emotion recognition</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>240</volume>:<fpage>122579</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.122579</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aslam</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sargano</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Habib</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Attention-based multimodal sentiment analysis and emotion recognition using deep neural networks</article-title>. <source>Appl Soft Comput</source>. <year>2023</year>;<volume>144</volume>:<fpage>110494</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.asoc.2023.110494</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Gated attention fusion network for multimodal sentiment classification</article-title>. <source>Knowl Based Syst</source>. <year>2022</year>;<volume>240</volume>:<fpage>108107</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2021.108107</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Learning transferable visual models from natural language supervision</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2021</year>. p. <fpage>8748</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nie</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Feature-attentive multimodal emotion analyzer with CLIP</article-title>. In: <conf-name>2023 International Conference on Image Processing, Computer Vision and Machine Learning (ICICML)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>317</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An image is worth 16 &#x00D7; 16 words: transformers for image recognition at scale</article-title>. <comment>arXiv:2010.11929. 2020</comment>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name></person-group>. <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. In: <conf-name>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</conf-name>. <publisher-loc>Kerrville, TX, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2019</year>. p. <fpage>4171</fpage>&#x2013;<lpage>86</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kanayama</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nasukawa</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Fully automatic lexicon expansion for domain-oriented sentiment analysis</article-title>. In: <conf-name>Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing</conf-name>. <publisher-loc>Kerrville, TX, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2006</year>. p. <fpage>355</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alwan</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Abd</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Sadiq</surname> <given-names>AT</given-names></string-name>, <string-name><surname>Khalaf</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liatsis</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Political Arabic articles orientation using rough set theory with sentiment lexicon</article-title>. <source>IEEE Access</source>. <year>2021</year>;<volume>9</volume>:<fpage>24475</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1109/access.2021.3054919</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Esuli</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sebastiani</surname> <given-names>F</given-names></string-name></person-group>. <article-title>SentiWordNet: a publicly available lexical resource for opinion mining</article-title>. In: <conf-name>International Conference on Language Resources and Evaluation (LREC 2006)</conf-name>. <publisher-loc>Paris, France</publisher-loc>: <publisher-name>ELRA</publisher-name>; <year>2006</year>. p. <fpage>417</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hamouda</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rohaim</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Reviews classification using sentiwordnet lexicon</article-title>. In: <conf-name>Proceedings of the World Congress on Computer Science and Information Technology (WCSIT)</conf-name>. <publisher-loc>Dubai, United Arab Emirates</publisher-loc>: <publisher-name>WCSIT</publisher-name>; <year>2011</year>. p. <fpage>104</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Naz</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Sentiment classification on twitter data using support vector machine</article-title>. In: <conf-name>2018 IEEE/WIC/ACM International Conference on Web Intelligence (WI)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>676</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Bayesian Na&#x00EF;ve Bayes classifiers to text classification</article-title>. <source>J Inform Sci</source>. <year>2018</year>;<volume>44</volume>(<issue>1</issue>):<fpage>48</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1177/0165551516677946</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Goel</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gautam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Real time sentiment analysis of tweets using Naive Bayes</article-title>. In: <conf-name>2016 2nd International Conference on Next Generation Computing Technologies (NGCT)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>257</fpage>&#x2013;<lpage>61</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rathor</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dimri</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Comparative study of machine learning approaches for Amazon reviews</article-title>. <source>Procedia Comput Sci</source>. <year>2018</year>;<volume>132</volume>:<fpage>1552</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.procs.2018.05.119</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Convolutional neural networks for sentence classification</article-title>. <comment>arXiv:1408.5882. 2014</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Attention-based LSTM network for cross-lingual sentiment classification</article-title>. In: <conf-name>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</conf-name>. <publisher-loc>Kerrville, TX, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2016</year>. p. <fpage>247</fpage>&#x2013;<lpage>56</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Tree-structured regional CNN-LSTM model for dimensional sentiment analysis</article-title>. <source>IEEE ACM Trans Audio Speech Lang Process</source>. <year>2019</year>;<volume>28</volume>:<fpage>581</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taslp.2019.2959251</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Siersdorfer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Minack</surname> <given-names>E</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hare</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Analyzing and predicting sentiment of images on the social web</article-title>. In: <conf-name>Proceedings of the 18th ACM International Conference on Multimedia</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2010</year>. p. <fpage>715</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Borth</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Breuel</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>SF</given-names></string-name></person-group>. <article-title>Large-scale visual sentiment ontology and detectors using adjective noun pairs</article-title>. In: <conf-name>Proceedings of the 21st ACM International Conference on Multimedia</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2013</year>. p. <fpage>223</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>How do your friends on social media disclose your emotions?</article-title> <source>Proc AAAI Conf Artif Intell</source>. <year>2014</year>;<volume>28</volume>(<issue>1</issue>):<fpage>306</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>K</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ling</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Mei</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Boosting image sentiment analysis with visual attention</article-title>. <source>Neurocomputing</source>. <year>2018</year>;<volume>312</volume>:<fpage>218</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2018.05.104</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Islam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Visual sentiment analysis for social images using transfer learning approach</article-title>. In: <conf-name>2016 IEEE International Conferences on Big Data and Cloud Computing (BDCloud), Social Computing and Networking (SocialCom), Sustainable Computing and Communications (SustainCom) (BDCloud-SocialCom-SustainCom)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>124</fpage>&#x2013;<lpage>30</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sermanet</surname> <given-names>P</given-names></string-name>, <string-name><surname>Reed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anguelov</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Going deeper with convolutions</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdullah</surname> <given-names>SMSA</given-names></string-name>, <string-name><surname>Ameen</surname> <given-names>SYA</given-names></string-name>, <string-name><surname>Sadeeq</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Zeebaree</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Multimodal emotion recognition using deep learning</article-title>. <source>J Appl Sci Technol Trends</source>. <year>2021</year>;<volume>2</volume>(<issue>1</issue>):<fpage>73</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.38094/jastt20291</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>You</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Cross-modality consistent regression for joint visual-textual sentiment analysis of social multimedia</article-title>. In: <conf-name>Proceedings of the Ninth ACM International Conference on Web Search and Data Mining</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2016</year>. p. <fpage>13</fpage>&#x2013;<lpage>22</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>MCH</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An image-text consistency driven multimodal sentiment analysis approach for social media</article-title>. <source>Inform Process Manage</source>. <year>2019</year>;<volume>56</volume>(<issue>6</issue>):<fpage>102097</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ipm.2019.102097</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Geng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Sentiment analysis of social media via multimodal feature fusion</article-title>. <source>Symmetry</source>. <year>2020</year>;<volume>12</volume>(<issue>12</issue>):<fpage>2010</fpage>. doi:<pub-id pub-id-type="doi">10.3390/sym12122010</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Emotion recognition from large-scale video clips with cross-attention and hybrid feature weighting neural networks</article-title>. <source>Int J Environ Res Public Health</source>. <year>2023</year>;<volume>20</volume>(<issue>2</issue>):<fpage>1400</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ijerph20021400</pub-id>; <pub-id pub-id-type="pmid">36674161</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Exploiting BERT for multimodal target sentiment classification through input space translation</article-title>. In: <conf-name>Proceedings of the 29th ACM International Conference on Multimedia</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2021</year>. p. <fpage>3034</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Unicoder-VL: a universal encoder for vision and language by cross-modal pre-training</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>(<issue>7</issue>):<fpage>11336</fpage>&#x2013;<lpage>44</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Clip-MSA: incorporating inter-modal dynamics and common knowledge to multimodal sentiment analysis with clip</article-title>. In: <conf-name>ICASSP 2024&#x2013;2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>8145</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>An</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wan Zainon</surname> <given-names>WMN</given-names></string-name></person-group>. <article-title>Improving multimodal sentiment prediction through vision-language feature interaction</article-title>. <source>Multimed Syst</source>. <year>2025</year>;<volume>31</volume>(<issue>1</issue>):<fpage>63</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s00530-024-01659-4</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>Attention is all you need</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Guyon</surname> <given-names>I</given-names></string-name>, <string-name><surname>Luxburg</surname> <given-names>UV</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wallach</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fergus</surname> <given-names>R</given-names></string-name>, <string-name><surname>Vishwanathan</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group>, editors. <source>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>.; <year>2017</year>. p. <fpage>6000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Carion</surname> <given-names>N</given-names></string-name>, <string-name><surname>Massa</surname> <given-names>F</given-names></string-name>, <string-name><surname>Synnaeve</surname> <given-names>G</given-names></string-name>, <string-name><surname>Usunier</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kirillov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zagoruyko</surname> <given-names>S</given-names></string-name></person-group>. <article-title>End-to-end object detection with transformers</article-title>. In: <conf-name>European Conference on Computer Vision</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2020</year>. p. <fpage>213</fpage>&#x2013;<lpage>29</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Niu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pang</surname> <given-names>L</given-names></string-name>, <string-name><surname>El Saddik</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Sentiment analysis on multi-view social data</article-title>. In: <conf-name>International Conference on Multimedia Modeling</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2016</year>. p. <fpage>15</fpage>&#x2013;<lpage>27</lpage>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>W</given-names></string-name></person-group>. <article-title>MultiSentiNet: a deep semantic network for multimodal sentiment analysis</article-title>. In: <conf-name>Proceedings of the 2017 ACM on Conference on Information and Knowledge Management</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2017</year>. p. <fpage>2399</fpage>&#x2013;<lpage>402</lpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A co-memory network for multimodal sentiment analysis</article-title>. In: <conf-name>The 41st International ACM SIGIR Conference on Research &#x0026; Development in Information Retrieval</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2018</year>. p. <fpage>929</fpage>&#x2013;<lpage>32</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Image-text sentiment analysis via deep multimodal attentive fusion</article-title>. <source>Knowl-Based Syst</source>. <year>2019</year>;<volume>167</volume>:<fpage>26</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2019.01.019</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Image-text multimodal emotion classification via multi-view attentional network</article-title>. <source>IEEE Trans Multimed</source>. <year>2020</year>;<volume>23</volume>:<fpage>4014</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2020.3035277</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>T</given-names></string-name></person-group>. <article-title>CLMLF: a contrastive learning and multi-layer fusion method for multimodal sentiment detection</article-title>. <comment>arXiv:2204.05515. 2022</comment>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Cross-modal sentiment analysis based on CLIP image-text attention interaction</article-title>. <source>Int J Adv Comput Sci Appl</source>. <year>2024</year>;<volume>15</volume>(<issue>2</issue>):<fpage>895</fpage>&#x2013;<lpage>903</lpage>. doi:<pub-id pub-id-type="doi">10.14569/ijacsa.2024.0150290</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multimodal sentiment analysis with image-text interaction network</article-title>. <source>IEEE Trans Multimed</source>. <year>2022</year>;<volume>25</volume>:<fpage>3375</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2022.3160060</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ping</surname> <given-names>W</given-names></string-name>, <string-name><surname>Molchanov</surname> <given-names>P</given-names></string-name>, <string-name><surname>Shoeybi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name></person-group>. <article-title>VILA: on pre-training for visual language models</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>26689</fpage>&#x2013;<lpage>99</lpage>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chay-intr</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Viriyayudhakorn</surname> <given-names>K</given-names></string-name>, <string-name><surname>Theeramunkong</surname> <given-names>T</given-names></string-name></person-group>. <article-title>LLaVAC: fine-tuning LLaVA as a multimodal sentiment classifier</article-title>. <comment>arXiv:2502.02938. 2025</comment>.</mixed-citation></ref>
</ref-list>
</back></article>