<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">71656</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.071656</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>GLAMSNet: A Gated-Linear Aspect-Aware Multimodal Sentiment Network with Alignment Supervision and External Knowledge Guidance</article-title>
<alt-title alt-title-type="left-running-head">GLAMSNet: A Gated-Linear Aspect-Aware Multimodal Sentiment Network with Alignment Supervision and External Knowledge Guidance</alt-title>
<alt-title alt-title-type="right-running-head">GLAMSNet: A Gated-Linear Aspect-Aware Multimodal Sentiment Network with Alignment Supervision and External Knowledge Guidance</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Dan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Zhoubin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Xia</surname><given-names>Yuze</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>yuzex@student.must.edu.mo</email></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yu</surname><given-names>Zhenhua</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>zhenhuayu@xust.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Artificial Intelligence &#x0026; Computer Science, Xi&#x2019;an University of Science and Technology</institution>, <addr-line>Xi&#x2019;an, 710054</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Institute of Systems Engineering, Macau University of Science and Technology</institution>, <addr-line>Macau, 999078</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Yuze Xia. Email: <email>yuzex@student.must.edu.mo</email>; Zhenhua Yu. Email: <email>zhenhuayu@xust.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>10</month><year>2025</year>
</pub-date>
<volume>85</volume>
<issue>3</issue>
<fpage>5823</fpage>
<lpage>5845</lpage>
<history>
<date date-type="received">
<day>09</day>
<month>08</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>09</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_71656.pdf"></self-uri>
<abstract>
<p>Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to detect sentiment polarity toward specific aspects by leveraging both textual and visual inputs. However, existing models suffer from weak aspect-image alignment, modality imbalance dominated by textual signals, and limited reasoning for implicit or ambiguous sentiments requiring external knowledge. To address these issues, we propose a unified framework named Gated-Linear Aspect-Aware Multimodal Sentiment Network (GLAMSNet). First of all, an input encoding module is employed to construct modality-specific and aspect-aware representations. Subsequently, we introduce an image&#x2013;aspect correlation matching module to provide hierarchical supervision for visual-textual alignment. Building upon these components, we further design a Gated-Linear Aspect-Aware Fusion (GLAF) module to enhance aspect-aware representation learning by adaptively filtering irrelevant textual information and refining semantic alignment under aspect guidance. Additionally, an External Language Model Knowledge-Guided mechanism is integrated to incorporate sentiment-aware prior knowledge from GPT-4o, enabling robust semantic reasoning especially under noisy or ambiguous inputs. Experimental studies conducted based on Twitter-15 and Twitter-17 datasets demonstrate that the proposed model outperforms most state-of-the-art methods, achieving 79.36% accuracy and 74.72% F1-score, and 74.31% accuracy and 72.01% F1-score, respectively.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Sentiment analysis</kwd>
<kwd>multimodal aspect-based sentiment analysis</kwd>
<kwd>cross-modal alignment</kwd>
<kwd>multimodal sentiment classification</kwd>
<kwd>large language model</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Nature Science Foundation of China</funding-source>
<award-id>62476216</award-id>
<award-id>62273272</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Key Research and Development Program of Shaanxi Province</funding-source>
<award-id>2024GX-YBXM-146</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Education Department of Shaanxi Provincial Government</funding-source>
<award-id>23JP091</award-id>
</award-group>
<award-group id="awg4">
<funding-source>Youth Innovation Team of Shaanxi Universities</funding-source>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Aspect-based Sentiment Analysis (ABSA) [<xref ref-type="bibr" rid="ref-1">1</xref>] has emerged as a critical subtask in sentiment analysis, and it provides fine-grained predictions by identifying the sentiment polarity toward specific aspect terms within a sentence. Unlike traditional sentence-level sentiment classification, which assigns a global label to an entire sentence or document, ABSA explicitly focuses on the sentiment expressed toward individual targets, such as product features, named entities, or events. Neural network models for ABSA, such as Attention-based LSTM with Aspect Embedding (ATAE-LSTM) [<xref ref-type="bibr" rid="ref-2">2</xref>] and Interactive Attention Network (IAN) [<xref ref-type="bibr" rid="ref-3">3</xref>], employ target-aware attention mechanisms to capture sentiment-bearing expressions relevant to a given aspect. While the early methods have significantly advanced fine-grained sentiment modeling, their exclusive dependence on textual data reduces performance in cases of ambiguous expressions, sarcasm, or contextually ambiguous language. Furthermore, pre-trained language models like Bidirectional Encoder Representations from Transformers (BERT) [<xref ref-type="bibr" rid="ref-4">4</xref>] and Bidirectional and Auto-Regressive Transformers (BART) [<xref ref-type="bibr" rid="ref-5">5</xref>] have enhanced contextual understanding in pure-text settings, it is still of great difficulty in resolving sentiment cues with the requirements of commonsense reasoning. Such constraints are especially critical in informal or multimodal settings [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>], where relying exclusively on textual data can compromise sentiment inference accuracy.</p>
<p>Multimodal Aspect-based Sentiment Analysis (MABSA) [<xref ref-type="bibr" rid="ref-8">8</xref>] becomes a promising direction with its capability in integrating both textual and visual modalities to enhance sentiment comprehension. This paradigm is of increasing relevance in dealing the information on platforms such as Twitter and Instagram, in which the users post comments frequently including images that convey rich emotional signals. MABSA is also considered as an important application scenario in the broader domain of affective computing [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>], where understanding sentiment across heterogeneous modalities is essential for developing emotionally intelligent AI systems. Multimodal Interactive Memory Network (MIMN) [<xref ref-type="bibr" rid="ref-11">11</xref>] and Target-oriented Multimodal BERT (TomBERT) [<xref ref-type="bibr" rid="ref-12">12</xref>] demonstrate that it can boost sentiment prediction performance by incorporating visual cues into aspect-aware textual representations, particularly in case of involving subtle emotional indicators or ambiguous expressions. Subsequent models, such as Entity Sensitive Attention and Fusion Network (ESAFN) [<xref ref-type="bibr" rid="ref-13">13</xref>], further advance cross-modal interaction by modeling target-image relevance through attention mechanisms or alignment strategies. Recent studies have incorporated external knowledge, such as commonsense or affective cues, to further enhance multimodal sentiment reasoning.</p>
<p>Building on these advances, several recent multimodal aspect-based sentiment analysis models have explored complementary strategies, yet several important gaps remain to be addressed. The Hierarchical Interactive Multimodal Transformer (HIMT) [<xref ref-type="bibr" rid="ref-14">14</xref>] strengthens image&#x2013;text interactions through hierarchical cross-modal attention; however, it does not impose explicit aspect&#x2013;region supervision, thus attention can drift to salient but aspect-irrelevant regions when visual evidence is weak or ambiguous. The Knowledge-Enhanced Fusion framework (KEF) [<xref ref-type="bibr" rid="ref-15">15</xref>] enriches textual semantics with external commonsense, but its fusion remains predominantly text-led, leaving modality imbalance unresolved when visual cues are informative. The global-local features fusion with co-attention (GLFFCA) [<xref ref-type="bibr" rid="ref-16">16</xref>] captures interactions at multiple scales, but its reliance on implicit attention weights limits interpretability and does not provide stable alignment signals under noisy inputs. Taken together, these observations indicate that robust aspect&#x2013;region alignment, balanced fusion, and stronger reasoning for ambiguous content remain open issues.</p>
<p>Therefore, we summarize three key challenges that remain unsolved in MABSA: (1) insufficient aspect&#x2013;region alignment, (2) modality imbalance in fusion, and (3) lack of semantic reasoning for vague or shifted inputs requiring external priors. These issues hinder performance in weak-signal or ambiguous scenarios. To address them, we propose GLAMSNet, integrating alignment supervision, aspect-aware gated fusion, and large language model-guided sentiment reasoning. The significance of this study lies on that GLAMSNet unifies hierarchical visual-textual alignment, gated aspect-aware fusion, and external language model guidance for precise, interpretable multimodal sentiment modeling. It effectively handles weak, conflicting, or noisy multimodal signals, substantially improving accuracy and robustness in real-world scenarios.</p>
<p>The major contributions of this paper are as follows:
<list list-type="simple">
<list-item><label>(1)</label>
<p>We introduce a hierarchical correspondence framework to align aspect terms with visual regions at coarse and fine granularities, enabling the model to focus on aspect-relevant regions, filter noise, and produce interpretable visualizations.</p></list-item>
<list-item><label>(2)</label>
<p>We propose the Gated-Linear Aspect-Aware Fusion (GLAF) module, which adaptively modulates aspect-aware textual features by semantic relevance, suppressing noisy or weakly aligned signals and enhancing fine-grained sentiment reasoning. The ReLU-based linear attention is extended to multimodal sentiment analysis for efficient aspect-aware cross-modal modeling.</p></list-item>
<list-item><label>(3)</label>
<p>We integrate external generative guidance from large language models (e.g., GPT-4o) to provide auxiliary sentiment predictions and embedding-level cues, improving inference under vague or noisy conditions and enhancing robustness on challenging multimodal inputs.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> reviews related works in text-based ABSA, multimodal sentiment analysis, and advances in knowledge integration and cross-modal alignment. <xref ref-type="sec" rid="s3">Section 3</xref> presents the GLAMSNet model with three core functions&#x2014;modality-specific representation with supervised alignment, aspect-aware fusion via a gated-linear mechanism, and incorporation of external sentiment knowledge for enhanced reasoning&#x2014;implemented through four modules. <xref ref-type="sec" rid="s4">Section 4</xref> reports experimental results, and <xref ref-type="sec" rid="s5">Section 5</xref> concludes with future research directions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Aspect-Based Sentiment Analysis</title>
<p>Aspect-based sentiment analysis (ABSA) [<xref ref-type="bibr" rid="ref-1">1</xref>] has been extensively explored within the textual modality, and existing methods primarily focus on modeling the interaction between aspect terms and their surrounding context. For instance, Recurrent Attention Memory Network (RAM) [<xref ref-type="bibr" rid="ref-17">17</xref>] employs a multi-hop memory mechanism to iteratively refine sentiment understanding, thereby improving robustness in complex sentence structures. Transformation Network (TNet) [<xref ref-type="bibr" rid="ref-18">18</xref>] introduces a context-preserving transformation mechanism to mitigate semantic drift during feature projection. Building upon these foundations, Multi-Grained Attention Network (MGAN) [<xref ref-type="bibr" rid="ref-19">19</xref>] proposes a hierarchical alignment strategy that aligns fine-grained aspect representations with coarse-grained sentence semantics.</p>
<p>While most research of ABSA concentrate on textual data, several studies explored vision-only sentiment classification. For example, Res-Aspect [<xref ref-type="bibr" rid="ref-20">20</xref>] introduces an aspect-aware visual attention mechanism to identify sentiment-relevant regions within an image based on a given aspect. By leveraging convolutional features and object-level cues, the model attempts to infer aspect-specific sentiment without textual input. However, vision-only methods face inherent limitations: visual semantics are frequently ambiguous, and the absence of textual context often leads to misinterpretation&#x2014;particularly in abstract or symbolic imagery. These challenges underscore the requirement for multimodal frameworks [<xref ref-type="bibr" rid="ref-21">21</xref>] that jointly leverage complementary information from both text and image for more accurate and robust aspect-level sentiment analysis.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Multimodal Aspect-Based Sentiment Analysis</title>
<p>To overcome limitations of unimodal approaches, multimodal aspect-based sentiment analysis (MABSA) [<xref ref-type="bibr" rid="ref-8">8</xref>] integrates text and image for richer sentiment understanding. Multimodal Interactive Memory Network (MIMN) [<xref ref-type="bibr" rid="ref-11">11</xref>] adopts dynamic memory for aspect-guided interaction, and Entity Sensitive Attention and Fusion Network (ESAFN) [<xref ref-type="bibr" rid="ref-13">13</xref>] uses selective attention to highlight relevant regions. Vision-and-Language BERT (ViLBERT) [<xref ref-type="bibr" rid="ref-22">22</xref>] enables co-attentional encoding of both modalities, while Target-oriented Multimodal BERT (TomBERT) [<xref ref-type="bibr" rid="ref-12">12</xref>] incorporates target-specific attention but lacks explicit alignment. Caption-Transformer BERT (CapTrBERT) [<xref ref-type="bibr" rid="ref-23">23</xref>] and the Recurrent Attention Network [<xref ref-type="bibr" rid="ref-24">24</xref>] further enhances fusion via auxiliary captions, saliency maps, or hierarchical interaction. However, most methods remain limited by static or text-biased fusion, reducing robustness under modality imbalance or implicit sentiment cues.</p>
<p>A considerable body of research has also followed the BERT-family combined with convolutional/ region-based vision backbones in MABSA. TomBERT [<xref ref-type="bibr" rid="ref-12">12</xref>] represents one of the earliest attempts, in which contextualized embeddings from BERT [<xref ref-type="bibr" rid="ref-4">4</xref>] are combined with global image features derived from a convolutional neural network, demonstrating the benefits of pre-trained transformers in capturing textual nuance. However, the reliance on global visual descriptors limited its ability to capture fine-grained sentiment cues. To address this issue, MIMN [<xref ref-type="bibr" rid="ref-11">11</xref>] subsequently introduced a multi-interactive memory network that jointly models textual and visual information under aspect supervision, using multi-hop memory updates to refine cross-modal representations. While effective, it relied on global ResNet features and thus lacks region-level alignment. Furthermore, ESAFN [<xref ref-type="bibr" rid="ref-13">13</xref>] incorporates explicit structure-aware attention to strengthen cross-modal interactions, ensuring that BERT-derived textual features are more directly matched with aspect-relevant image regions.</p>
<p>Recent models focus on finer-grained multimodal interaction. HIMT [<xref ref-type="bibr" rid="ref-14">14</xref>] enhances multi-level semantic alignment between aspect terms and visual content, while KEF [<xref ref-type="bibr" rid="ref-15">15</xref>] incorporates commonsense knowledge to support sentiment reasoning. Within the same BERT &#x002B; CNN lineage, KEF employs knowledge-enhanced fusion to refine semantic correspondence between aspect terms and visual objects, while HIMT develops hierarchical alignment strategies that progressively link coarse-grained and fine-grained features across modalities. These advances collectively demonstrate that BERT-family text encoders, when paired with region-level visual representations, form a powerful yet flexible backbone for multimodal sentiment classification. Nonetheless, existing methods often rely on co-attention or memory units alone, which struggle to fully resolve modality imbalance and ambiguities in implicit sentiment expressions.</p>
<p>Overall, the BERT &#x002B; CNN paradigm has become a central stream in multimodal aspect-based sentiment analysis. By combining contextualized textual embeddings from BERT-family encoders with convolutional or region-level visual features, these models demonstrate clear advantages in aligning aspect-aware semantics with visual evidence. Nevertheless, persistent challenges remain within this stream: early works such as TomBERT [<xref ref-type="bibr" rid="ref-12">12</xref>] depend on global CNN features and are unable to capture fine-grained cues; MIMN [<xref ref-type="bibr" rid="ref-11">11</xref>] introduce memory-based cross-modal interaction but is limited by its reliance on ResNet-level global representations; and more recent frameworks such as ESAFN [<xref ref-type="bibr" rid="ref-13">13</xref>], HIMT [<xref ref-type="bibr" rid="ref-14">14</xref>], and KEF [<xref ref-type="bibr" rid="ref-15">15</xref>] enhance alignment or incorporate external knowledge, yet they still struggled to overcome modality imbalance and implicit sentiment ambiguity. Taken together, this lineage underscores both the promise and the limitations of BERT &#x002B; CNN approaches, stressing the requirement for unified solutions that can effectively tackle the challenges of supervised alignment, adaptive fusion, and knowledge-guided reasoning.</p>
<p>To enhance interpretability and target-aware reasoning, Face-Sensitive Image-to-Emotional-Text Translation (FITE) [<xref ref-type="bibr" rid="ref-25">25</xref>] aligns facial expressions with emotional text for cross-modal grounding, while the Affective Region Recognition and Fusion Network (ARFN) [<xref ref-type="bibr" rid="ref-26">26</xref>] applies region-level filtering to improve robustness under visual noise. The Global&#x2013;Local Feature Fusion Network with Co-Attention (GLFFCA) [<xref ref-type="bibr" rid="ref-16">16</xref>] leverages gated fusion and contrastive learning for image&#x2013;text alignment, yet remains largely aspect-agnostic. Multimodal Dual Cause Analysis (MDCA) [<xref ref-type="bibr" rid="ref-27">27</xref>] introduces causal reasoning via vision&#x2013;language pretraining but lacks mechanisms to explicitly inject LLM-based sentiment predictions into the decision process. Most recently, the Prompt-guided Dual Query and Span-based Aspect Modeling (DQPSA) framework [<xref ref-type="bibr" rid="ref-28">28</xref>] enhances target-specific alignment by combining prompt-driven visual grounding with energy-based span prediction, offering a unified and task-aware solution to multimodal aspect-based sentiment analysis. Efficient Multimodal Transformer with Dual-Level Feature Restoration (EMT-DLFR) [<xref ref-type="bibr" rid="ref-29">29</xref>] and Modality-Invariant Contrastive Learning (MICL) [<xref ref-type="bibr" rid="ref-30">30</xref>] improve robustness under modality incompleteness, core challenges in alignment, fusion, and reasoning remain.</p>
<p>Beyond sentiment classification, related research in multimodal recommender systems has also demonstrated the benefits of combining affective cues from text and images. For instance, a study to explore how visual sentiment embedded in product images, together with textual sentiment in user reviews, can be fused within a context-aware collaborative filtering framework to enhance recommendation accuracy [<xref ref-type="bibr" rid="ref-31">31</xref>]. Although the objective differs from aspect-based sentiment classification, both domains share the insight that textual and visual modalities provide complementary emotional information. This cross-domain evidence further underscores the importance of multimodal fusion and motivates our work, which extends the paradigm by focusing on fine-grained aspect&#x2013;region alignment, adaptive fusion, and external knowledge guidance for robust sentiment reasoning.</p>
<p>Despite these advancements, current MABSA [<xref ref-type="bibr" rid="ref-8">8</xref>] models still suffer from three recurrent limitations: (1) the absence of supervised alignment between aspect terms and visual regions, (2) insufficient modeling of aspect-aware dynamic fusion, and (3) underutilization of external reasoning capabilities, especially those enabled by modern LLMs. These challenges highlight the need for a more unified and interpretable framework that can simultaneously address cross-modal alignment, fusion adaptability, and knowledge-guided reasoning.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>In this study, we propose a unified framework named Gated-Linear Aspect-Aware Multimodal Sentiment Network (GLAMSNet), as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>:</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>An overview of the GLAMSNet architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_71656-fig-1.tif"/>
</fig>
<p>The proposed framework comprises four modules: (1) Input Encoding and Modality Representation Construction, (2) Image&#x2013;Aspect Correlation Matching, (3) Gated-Linear Aspect-Aware Fusion (GLAF), and (4) External Language Model Knowledge-Guided. The fused representation is then fed into a sentiment classification layer for aspect-level polarity prediction. The first two modules (<xref ref-type="sec" rid="s3_2">Sections 3.2</xref> and <xref ref-type="sec" rid="s3_3">3.3</xref>) address insufficient aspect&#x2013;region alignment by constructing aspect-aware unimodal features and applying hierarchical alignment supervision. The GLAF module (<xref ref-type="sec" rid="s3_4">Section 3.4</xref>) suppresses irrelevant textual information and refines features based on aspect semantics, while the External Knowledge-Guided (<xref ref-type="sec" rid="s3_5">Section 3.5</xref>) incorporates sentiment priors from large language models to enhance reasoning under complex inputs.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Task Definition</title>
<p>Given a text sequence of <italic>n</italic> words <italic>S</italic> &#x003D; (<italic>w</italic><sub>1</sub>, <italic>w</italic><sub>2</sub>, &#x2026; , <italic>w</italic><sub><italic>n</italic></sub>) and its associated image <italic>V</italic>, as well as a specific aspect <italic>A</italic> &#x003D; (<italic>a</italic><sub>1</sub>, <italic>a</italic><sub>2</sub>, &#x2026;, <italic>a</italic><sub><italic>m</italic></sub>), <italic>m</italic> implies the length of aspesct phrase. Particularly, we assume that all aspect entities (i.e., words or phrases) in <italic>S</italic> have been provided. The specific aspect is assigned a label <italic>y</italic> &#x2208; {<italic>negative, neutral, positive</italic>}. Formally, given a training dataset comprising <italic>D</italic> instances, each sample includes a triplet <italic>X</italic> &#x003D; (<italic>S</italic>, <italic>V</italic>, <italic>A</italic>), where the model is expected to infer the sentiment polarity corresponding to the specified aspect term within the multimodal context. The objective is to learn a function capable of predicting aspect-level sentiment for multimodal samples.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Input Encoding and Modality Representation Construction</title>
<p>Let us consider a sample set with each sample represented as a triplet (<italic>S</italic>, <italic>V</italic>, <italic>A</italic>), which consists of a textual input, a paired image, and a set of aspect terms. In this module, we independently encode the textual and visual modalities to obtain modality-specific representations that serve as the input for downstream reasoning modules.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Backbone Choice</title>
<p>In the proposed framework, we adopt RoBERTa [<xref ref-type="bibr" rid="ref-32">32</xref>] as the textual encoder and Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] as the visual backbone. This choice is motivated by following the requirements of multimodal aspect-based sentiment analysis. The textual inputs in the applied datasets (e.g., Twitter-15 and Twitter-17 datasets) are typically short, noisy, and lack grammatical regularity, which demands a backbone capable of stable representation learning under limited data conditions.</p>
<p>On the textual side, RoBERTa [<xref ref-type="bibr" rid="ref-32">32</xref>], a robustly optimized variant of BERT [<xref ref-type="bibr" rid="ref-4">4</xref>], has demonstrated strong contextual modeling and reliable fine-tuning in low-resource sentiment tasks. By contrast, Decoding-enhanced BERT with Disentangled Attention (DeBERTa) [<xref ref-type="bibr" rid="ref-34">34</xref>] achieves gains on large-scale benchmarks but usually requires abundant data to exploit its disentangled attention fully. When applied to small and noisy corpora such as Twitter, DeBERTa [<xref ref-type="bibr" rid="ref-34">34</xref>] often suffers from unstable convergence and overfitting, whereas RoBERTa [<xref ref-type="bibr" rid="ref-32">32</xref>] provides proven robustness.</p>
<p>On the visual side, sentiment cues are often localized in specific regions, such as faces, trophies, or symbolic objects, rather than spread across the entire image. Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] is particularly suitable for this scenario because its region proposal mechanism yields object-level features that align naturally with aspect&#x2013;image correspondence. In contrast, patch-based encoders such as the Vision Transformer (ViT) [<xref ref-type="bibr" rid="ref-35">35</xref>] or holistic vision&#x2013;language models such as Contrastive Language&#x2013;Image Pre-training (CLIP) [<xref ref-type="bibr" rid="ref-36">36</xref>] and Bootstrapping Language&#x2013;Image Pre-training (BLIP) [<xref ref-type="bibr" rid="ref-37">37</xref>] primarily produce global embeddings that dilute fine-grained sentiment cues, thereby limiting their effectiveness in aspect-level alignment.</p>
<p>Another key consideration is efficiency and compatibility. Pre-extracted region proposals from Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] enable efficient training on a single GPU, avoiding the latency and memory overhead of large vision&#x2013;language frameworks. More importantly, the discrete region-level features extracted by Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] and the stable token representations produced by RoBERTa [<xref ref-type="bibr" rid="ref-32">32</xref>] provide complementary granularity that aligns naturally with our Gated-Linear Aspect-Aware Fusion (GLAF) and alignment objectives, both of which require fine-grained and structurally consistent inputs to achieve effective cross-modal integration. In contrast, global embeddings generated by models such as ViT [<xref ref-type="bibr" rid="ref-35">35</xref>], CLIP [<xref ref-type="bibr" rid="ref-36">36</xref>], or BLIP [<xref ref-type="bibr" rid="ref-37">37</xref>] are less amenable to alignment-based supervision, while DeBERTa&#x2019;s [<xref ref-type="bibr" rid="ref-34">34</xref>] instability on small and noisy corpora undermines downstream fusion. Therefore, the combination of RoBERTa [<xref ref-type="bibr" rid="ref-32">32</xref>] and Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] provides both efficiency and natural compatibility with GLAF and alignment-driven optimization, making it a well-suited backbone for our framework.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Textual Modality Encoder</title>
<p>We utilize a pretrained RoBERTa language model [<xref ref-type="bibr" rid="ref-32">32</xref>] to encode the input sentence in an aspect-aware manner. For the original input sentence <italic>S</italic> and an aspect term <italic>a</italic><sub><italic>m</italic></sub> &#x2208; <italic>A</italic>, a modified sequence <italic>S</italic><sup>&#x2032;</sup> is constructed by concatenating the aspect term with a masked version of the sentence, separated by a special delimiter:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Concat</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>P</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>S</italic><sub><italic>masked</italic></sub> is derived by replacing the aspect span <italic>a</italic><sub><italic>m</italic></sub> in <italic>S</italic> with a placeholder token. The token [<italic>SEP</italic>] is a special separator symbol used in RoBERTa to mark boundaries between different segments. The function Concat (&#x00B7;) denotes the token-level sequence concatenation (not embedding concatenation), and the square brackets [&#x00B7;] indicate token sequences rather than scalar values or mathematical sets.</p>
<p>The resulting sequence <italic>S</italic><sup>&#x2032;</sup> is then encoded to produce contextualized token embeddings:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>RoBERTa</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <italic>H</italic><sub><italic>T</italic></sub> serves as the aspect-aware textual embedding for subsequent cross-modal reasoning modules, <italic>h</italic> is the hidden dimension, and <italic>l</italic> is the input length. All visual and textual feature matrices are organized in a column-wise format, and each column represents a single token or region embedding vector of dimension &#x210E;.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Visual Modality Encoder</title>
<p>To extract region-level visual features relevant to the aspect, a pretrained object detection model named Faster R-CNN [<xref ref-type="bibr" rid="ref-33">33</xref>] is used to process the input image <italic>V</italic>. The detector identifies the top <italic>K</italic> regions with the highest confidence scores, Subsequently, the detected object proposals are ranked based on their category-wise confidence scores, and the top 100 proposals are retained to preserve finer-grained objects, thereby facilitating precise alignment with aspect targets. Let <italic>R, R</italic> &#x2208; <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup>2048&#x00D7;100</sup> denotes the regional representations, we have:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Faster</mml:mtext></mml:mrow><mml:mrow><mml:mtext>R</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>CNN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>V</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In order to capture the inter-object dependencies, the regional feature matrix <italic>R</italic> is projected and subsequently input into a Transformer encoder, producing object-level visual representations:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Transformer</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:mrow></mml:msubsup><mml:mi>R</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>W</italic><sub><italic>R</italic></sub> &#x2208; <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup>2048&#x00D7;<italic>h</italic></sup> and <italic>H</italic><sub><italic>V</italic></sub> &#x2208; <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;100</sup>.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Image&#x2013;Aspect Correlation Matching</title>
<p>In multimodal aspect-based sentiment analysis, not all images are semantically aligned with the aspect under discussion in the text. In other words, some visual content may be unrelated or even with semantic noise, thereby degrading classification performance if they are fused indiscriminately. In this study, an image&#x2013;aspect relevance module [<xref ref-type="bibr" rid="ref-38">38</xref>] is introduced to evaluate the semantic relatedness between the image and a given aspect. This module provides two levels of alignment supervision&#x2014;coarse-grained image-aspect relevance and fine-grained region-level attention refinement&#x2014;to guide the model in suppressing irrelevant visual content and attending to sentiment-relevant regions.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Coarse-Grained Image-Aspect Relevance</title>
<p>Let us start with the construction of visual representations that are guided by the aspect-aware textual semantics. We adopt a Cross-Modal Transformer layer [<xref ref-type="bibr" rid="ref-39">39</xref>] to model the interaction between the aspect and the image, which regards image representation <italic>H</italic><sub><italic>V</italic></sub> as queries, and contextualized aspect representations <italic>H</italic><sub><italic>T</italic></sub> as keys and values as follows:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>CrossModal-Transformer</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> &#x2208; <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;100</sup> is the generated aspect-based image representation. This cross-modal attention allows each image region to query the aspect-aware textual features and retrieve semantically aligned context, thereby generating visual representations that are conditioned on the corresponding aspect term.</p>
<p>To assess whether the image is semantically relevant to the aspect, we apply max-pooling to the visual representation <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to extract the most salient features for relevance prediction. This downsampling operation selects the maximum activation across spatial dimensions for each feature channel, highlighting informative regions while suppressing noise:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>MaxPool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The aggregated vector <italic>h</italic><sub><italic>v</italic></sub> is subsequently fed into a binary classifier to compute the relevance score <italic>P</italic> (<italic>r</italic>) &#x2208; (0, 1), which represents the probability that the image is semantically aligned with the current aspect:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>W</italic><sub><italic>r</italic></sub> &#x2208; <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup>1&#x00D7;<italic>h</italic></sup> and <italic>b</italic><sub><italic>r</italic></sub> &#x2208; <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula> are learnable parameters of the classifier, and <italic>&#x03C3;</italic> (&#x00B7;) denotes the sigmoid activation function.</p>
<p>A high relevance score indicates that the visual content provides meaningful context for the given aspect, while a low score suggests semantic misalignment. This scalar value <italic>P</italic> (<italic>r</italic>) serves as the global relevance signal and would be utilized in subsequent gating to suppress noisy or unrelated visual information during fusion.</p>
<p>While cross-modal fusion enables rich semantic interactions, irrelevant or noisy visual information may still introduce confusion if indiscriminately combined with textual representations. we leverage the scalar relevance score <italic>P</italic> (<italic>r</italic>) &#x2208; (0, 1) to suppress uninformative or misleading visual content. The score is then broadcast and expanded to match the shape of the visual representation <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, producing a gating matrix <italic>G</italic> &#x2208; <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;100</sup>, and each entry in <italic>G</italic> equals to <italic>P</italic> (<italic>r</italic>). The relevance-guided visual features are computed via element-wise multiplication:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mi>H</mml:mi><mml:mi>V</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mo>&#x2299;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>r</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula>where &#x2299; denotes the Hadamard (element-wise) product.</p>
<p>When <italic>P</italic> (<italic>r</italic>) is approaximate to 0 (indicating low image-aspect relevance), the visual signal is effectively suppressed. Conversely, when <italic>P</italic> (<italic>r</italic>) approaches to 1, the model retains full visual representation. This gated modulation acts as a soft attention filter, allowing the model to downweight noisy or irrelevant visual cues while preserving informative regions for downstream fusion.</p>
<p>To estimate semantic relevance between an image and a given aspect, we follow the coarse-grained region-aspect alignment [<xref ref-type="bibr" rid="ref-38">38</xref>], using manually annotated relevance tags to supervise visual&#x2013;textual correspondence. The predicted relevance score <italic>P</italic> (<italic>r</italic>) is optimized with binary cross-entropy loss:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>M</mml:mi></mml:mfrac><mml:mover><mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:mo stretchy="false">[</mml:mo></mml:mrow><mml:mi>M</mml:mi></mml:mover><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>r</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:math></disp-formula>where <italic>M</italic> denotes the number of training samples with relevance annotations; <italic>k</italic> denote the <italic>k</italic>-th training instance. A higher penalty is imposed when the model predicts high relevance scores for semantically unrelated image&#x2013;aspect pairs, or low scores for relevant pairs. This loss supervision guides the model to distinguish informative visual signals from noisy or irrelevant content.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Fine-Grained Region-Level Attention Refinement</title>
<p>While the visual modality provides abundant contextual signals for multimodal sentiment analysis, the sentiment expressed toward a given aspect is usually grounded in specific regions of the image rather than the entire visual scene. To be specific, for an individual mentioned in the text, the associated image may contain unrelated elements such as background objects, surrounding people, or environmental noise.</p>
<p>In this part, we employ a region alignment supervision module [<xref ref-type="bibr" rid="ref-38">38</xref>] to compute attention over localized image regions and optimize it toward a ground-truth alignment distribution based on manually annotated aspect-relevant regions. Through this supervision, the model is encouraged to focus on semantically relevant regions while suppressing noisy or unrelated visual content, thereby enhancing aspect-aware representation learning.</p>
<p>Cross-Modal Transformer layer [<xref ref-type="bibr" rid="ref-39">39</xref>] is applied to obtain the aspect-aware attention distribution over 100 object proposals from Faster R-CNN. We use the representation of the first token in the aspect input <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> as queries, and the filtered image representations <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msubsup><mml:mi>H</mml:mi><mml:mi>V</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> as keys and values:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>CrossModal-Transformer</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>V</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mi>V</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> &#x2208; <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;1</sup> is the generated image-based aspect representation. This aspect-aware visual representation <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> serves as both the region-level visual cue for multimodal fusion and a supervision target for alignment training. It enables the model to focus on sentiment-relevant regions in a fine-grained and interpretable manner.</p>
<p>We denote the attention weights from the <italic>i</italic>-th head of the CrossModal Transformer as <bold><italic>D</italic></bold><sub><italic>i</italic></sub>. To obtain the final attention distribution across the <italic>K</italic> candidate regions, we compute the mean of all <italic>m</italic> attention heads, resulting in <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi mathvariant="bold-italic">D</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>m</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <bold><italic>D</italic></bold> &#x2208; <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>K</italic></sup>.</p>
<p>To supervise the attention mechanism in learning semantically meaningful region-aspect alignments, we leverage the ground-truth (GT) bounding boxes that is manually annotated by human experts [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>]. These GT regions serve as soft alignment references, enabling the model to calibrate its attention distribution toward sentiment-relevant visual areas. For each training instance, we construct the alignment distribution based on the degree of spatial overlap between the detected candidate regions and the annotated GT box. We construct the supervised alignment distribution <italic>&#x03B1;</italic>&#x2208;<italic>K</italic> by computing the Intersection over Union (IoU) between each of the <italic>K</italic> &#x003D; 100 candidate regions&#x2014;extracted from image <italic>V</italic> using Faster R-CNN&#x2014;and the ground-truth (GT) bounding box <italic>B</italic><sub><italic>GT</italic></sub>. For the <italic>i</italic>-th region <italic>B</italic><sub><italic>i</italic></sub>, we calculate its IoU score with <italic>B</italic><sub><italic>GT</italic></sub>, denoted as <italic>s</italic><sub><italic>i</italic></sub> &#x003D; IoU(<italic>B</italic><sub><italic>i</italic></sub>, <italic>B</italic><sub><italic>GT</italic></sub>). If <italic>s</italic><sub><italic>i</italic></sub> exceeds a predefined threshold <italic>&#x03C4;</italic> &#x003D; 0.5, it is retained; otherwise, it is set to zero. The generated alignment distribution <italic>&#x03B1;</italic> serves as a soft supervision signal to guide the model&#x2019;s region-level attention mechanism toward aspect-relevant visual regions.</p>
<p>To strengthen the alignment between the model-generated attention distribution <bold><italic>D</italic></bold> and the human-annotated alignment distribution <italic>&#x03B1;</italic>, we adopt the Kullback&#x2013;Leibler (KL) divergence as the training objective. The region-level alignment loss is formally defined as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>ATT</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="bold-italic">D</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>N</italic> is the number of aspect&#x2013;image samples for region-level alignment training, we use manually annotated supervision with global image&#x2013;aspect relevance labels and fine-grained bounding boxes to guide attention learning. These annotations help the model distinguish sentiment-relevant regions from noise, providing structured guidance for cross-modal alignment.</p>
<p>However, while the alignment module improves aspect&#x2013;region correspondence, it lacks a fine-grained fusion of textual and image-derived aspect features. To address this, we introduce the Gated-Linear Aspect-Aware Fusion (GLAF) module, which adaptively integrates multimodal cues based on aspect relevance.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Gated-Linear Aspect-Aware Fusion</title>
<p>To improve aspect-aware representations and suppress noisy textual tokens, we propose the Gated-Linear Aspect-Aware Fusion (GLAF) module. Unlike standard cross-modal fusion, GLAF applies aspect-guided filtering within the textual modality, followed by lightweight attention-based refinement, enabling distilled sentiment-relevant signals and strengthened aspect-aware dependencies, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. It operates in two stages: the Adaptive Fusion Stage uses gating to modulate textual token contributions based on the aspect vector, enhancing relevant features; the Semantic Interaction Stage refines the fused representation through aspect-aware cross-modal interactions, capturing nuanced sentiment cues from both modalities.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Overview of the gated-linear aspect-aware fusion</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_71656-fig-2.tif"/>
</fig>
<p>Before describing the details of GLAF, we clarify the notations used in this section to avoid ambiguity. Specifically, <italic>H</italic><sub><italic>T</italic></sub> denotes the aspect-aware textual embedding obtained from the RoBERTa encoder (<xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>), while <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the image-based aspect representation derived from the Fine-grained region-level attention refinement module (<xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>). Thus, <italic>H</italic><sub><italic>T</italic></sub> carries purely textual signals, whereas <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> incorporates cross-modal cues guided by visual evidence. In the following subsections, these two representations are combined to perform gated fusion.</p>
<p>The rationale for using the image-based aspect representation <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to gate the textual embedding <italic>H</italic><sub><italic>T</italic></sub> lies in the complementary role of visual context. Social media text often contains ambiguous or noisy tokens (e.g., sarcasm, irrelevant hashtags), which are difficult to filter using text alone. The visual modality, however, provides additional grounding cues: if an aspect appears visually salient in the associated image, the corresponding textual tokens should be emphasized; otherwise, their contribution should be suppressed. Therefore, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> serves as an auxiliary signal that adaptively highlights sentiment-relevant tokens in <italic>H</italic><sub><italic>T</italic></sub>, while diminishing irrelevant textual content. In this way, the gating is not inverted but reflects the guiding role of visual evidence in resolving textual ambiguity.</p>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Adaptive Fusion Stage</title>
<p>To dynamically regulate the contribution of different textual features under aspect guidance, we develop a gating-based adaptive fusion strategy. Specifically, given the aspect-aware textual representation <italic>H</italic><sub><italic>T</italic></sub> and the generated image-based aspect representation <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, we first compute the fusion gating matrix <italic>G</italic> as:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>W</italic><sub>1</sub>, <italic>W</italic><sub>2</sub> &#x2208; <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;<italic>h</italic></sup> are learnable transformation matrices; <italic>b</italic> &#x2208; <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic></sup> is the bias term; and <italic>&#x03C3;</italic> (&#x00B7;) denotes the element-wise Sigmoid activation function.</p>
<p>The gating coefficient <italic>G</italic> &#x2208; (0, 1) indicates the contribution degree of the visual modality to the fusion representation. This gating formulation is a widely adopted linear gating strategy commonly used in multimodal representation learning [<xref ref-type="bibr" rid="ref-16">16</xref>]. Based on the gating matrix <italic>G</italic>, the fused representation is computed as:
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>G</mml:mi><mml:mo>&#x2299;</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>G</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>This fusion mechanism allows the model to automatically adjust the relative contributions of individual textual tokens according to their aspect relevance, effectively mitigating the impact of noisy or irrelevant textual signals.</p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Semantic Interaction Stage</title>
<p>Although the gating-based fusion mechanism adaptively regulates token contributions based on aspect relevance, it does not explicitly model fine-grained semantic dependencies across the fused sequence. To address this, we propose a lightweight semantic interaction module based on multi-head linear attention, building upon ReLU-based linear attention [<xref ref-type="bibr" rid="ref-42">42</xref>] with structural simplification and reduced computational overhead. Specifically, we remove the softmax in self-attention and replace it with ReLU-activated queries and keys for more efficient, stable computation. This design reduces numerical instability and strengthens cross-modal semantic interaction without overfitting to dominant features, while a residual connection preserves original fused features for fine-grained refinement guided by learned attention weights.</p>
<p>Let <italic>H</italic><sub><italic>F</italic></sub> denote the fused representation obtained from the previous stage, we project <italic>H</italic><sub><italic>F</italic></sub> into the query, key, and value spaces as follows:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ReLU</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <italic>W</italic><sub><italic>Q</italic></sub>, <italic>W</italic><sub><italic>K</italic></sub>, <italic>W</italic><sub><italic>V</italic></sub> &#x2208; <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;<italic>h</italic></sup> are learnable projection matrices.</p>
<p>The ReLU activation introduces non-linearity and enhances the representation capacity. These projected matrices are then fed into a linear attention layer to compute contextualized embeddings, which are subsequently used for sentiment classification.</p>
<p>Next, the attention weights are computed using a linear normalization strategy. Specifically, we define the unnormalized attention matrix as <italic>QK</italic><sup><italic>T</italic></sup>. The normalized attention weights are calculated as:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>Attn</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>K</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>K</mml:mi><mml:mi>j</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <italic>Q</italic><sub><italic>i</italic></sub> &#x2208; <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic></sup> and <italic>K</italic><sub><italic>j</italic></sub> &#x2208; <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic></sup> denote the <italic>i</italic>-th query and <italic>j</italic>-th key vector, respectively, and <italic>&#x03B5;</italic> is a small constant (e.g., 10<sup>&#x2212;6</sup>) added to prevent numerical instability.</p>
<p>The attention-weighted representation is computed as:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext>Attn</mml:mtext></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>V</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> &#x2208; <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic>&#x00D7;<italic>l</italic></sup> denotes the final enhanced fused representation incorporating contextual information via semantic interaction and <italic>W</italic><sub><italic>o</italic></sub> is a learnable projection matrix used as a linear transformation function. The residual connection with <italic>H</italic><sub><italic>F</italic></sub> helps preserve the original modality-aware features while injecting interaction-aware refinements.</p>
<p>This semantic interaction mechanism enables to model fine-grained global dependencies across tokens within the fused representation <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, thereby enhancing the precision and contextual richness of the aspect-level semantic encoding.</p>
<p>The fused representation is fed into a sentiment classification layer. For aspect-level prediction, we apply mean pooling over token-level representations, a common method in sentence representation learning [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>] for effectively aggregating embeddings.
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mtext>Pooling</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>The resulting vector <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is then fed into a linear classifier to predict the sentiment polarity distribution:
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>where <italic>P</italic> (<italic>y</italic>&#x2502;<italic>X</italic>) &#x2208; <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup>3</sup> denotes the probabilities for three sentiment classes: positive, neutral, and negative.</p>
</sec>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>External Language Model Knowledge-Guided</title>
<p>Although multimodal models can infer aspect-level sentiment from joint text&#x2013;image representations, predictions often suffer from instability due to limited training data and semantic inconsistencies, especially with low-quality images or ambiguous text. While GLAF mitigates modality imbalance and improves fusion, it still relies only on observed inputs and lacks deeper reasoning under weak or conflicting signals. To address this, we introduce an external language model knowledge-guided mechanism that uses large-scale generative models to provide complementary sentiment priors, enhancing training supervision, representation learning, and semantic alignment for better generalization.</p>
<sec id="s3_5_1">
<label>3.5.1</label>
<title>Sentiment Label Prediction and Embedding</title>
<p>To incorporate external sentiment supervision at the aspect level, we reformulate the original dataset into an aspect-specific instance structure. While each original sample is defined as a triplet <italic>X</italic> &#x003D; (<italic>S</italic>, <italic>V</italic>, <italic>A</italic>). Specifically, we query the LLM to infer a sentiment polarity label <italic>y</italic><sup><italic>GPT</italic></sup> &#x2208; {0, 1, 2} for each training sample based solely on its textual input. These predicted labels are then mapped into a continuous vector space via a learnable embedding function Emb: &#x2124; &#x2192; <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic></sup>, yielding sentiment-guided embedding vectors: <italic>e</italic><sub><italic>g</italic></sub> &#x003D; Emb (<italic>y</italic><sup><italic>GPT</italic></sup>), and <italic>e</italic><sub><italic>g</italic></sub> &#x2208; <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mrow><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula><sup><italic>h</italic></sup> represents the knowledge-guided sentiment prior for the <italic>g</italic>-th sample. The embedding matrix is jointly optimized with the rest of the network parameters during training, allowing the model to learn how to incorporate external sentiment knowledge effectively.</p>
</sec>
<sec id="s3_5_2">
<label>3.5.2</label>
<title>External Sentiment Label Generation</title>
<p>To obtain high-quality sentiment supervision signals, we leverage the reasoning capabilities of a pretrained large language model (GPT-4o) to generate auxiliary sentiment labels for each training instance. For a given sample <italic>X</italic> &#x003D; (<italic>S</italic>, <italic>V</italic>, <italic>A</italic>), a carefully designed prompt is constructed and fed into GPT-4o to infer a sentiment polarity label: <italic>y</italic><sup><italic>GPT</italic></sup> &#x2208; {negative, neutral, positive}.</p>
<p>Inspired by the prompt-based causal reasoning design in Multimodal Dual Cause Analysis (MDCA) [<xref ref-type="bibr" rid="ref-27">27</xref>], we design a set of structured prompt templates to elicit aspect-level sentiment understanding from the LLM.</p>
<p>The prompt templates include a system prompt, defining the LLM&#x2019;s role as a multimodal sentiment analysis expert, and a task prompt, providing text and visual context for aspect-level sentiment prediction. <xref ref-type="table" rid="table-1">Table 1</xref> shows representative examples illustrating their structural diversity and task relevance.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Prompt templates used for external sentiment label generation via GPT-4o</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Prompt type</th>
<th align="center">Prompt content</th>
</tr>
</thead>
<tbody>
<tr>
<td>System Prompt</td>
<td>You are a multimodal sentiment analysis expert, specializing in fine-grained sentiment prediction by integrating both textual and visual information.</td>
</tr>
<tr>
<td>Task Prompt</td>
<td>Given text <italic>S</italic>, image <italic>V</italic>, and aspect term <italic>A</italic>, predict the sentiment polarity (positive, neutral, negative) toward the aspect. Output in the format: &#x201C;Sentiment: [label]&#x201D;.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For each training sample, the overall procedure for external sentiment label generation via GPT-4o is formally in Algorithm 1.</p>
<fig id="fig-5">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_71656-fig-5.tif"/>
</fig>
<p>Specifically, a combined system and task prompt is constructed based on the input text, image, and aspect term, and subsequently fed into GPT-4o to generate a predicted sentiment label. The resulting label is stored in an external sentiment label repository, which serves as prior supervision for downstream knowledge-guided training.</p>
</sec>
<sec id="s3_5_3">
<label>3.5.3</label>
<title>Sentiment Representation Alignment Loss and Joint Optimization</title>
<p>In the multimodal fusion stage, the Gated-Linear Aspect-Aware Fusion (GLAF) module generates a fused sentiment representation <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, which integrates textual and visual features through gated fusion and linear attention mechanisms. To improve the alignment between the model&#x2019;s sentiment predictions and external knowledge, we introduce a cosine embedding loss that encourages semantic consistency between <italic>H&#x2019; F</italic> and the sentiment-guided vector <italic>e</italic>. The loss is formally defined as:
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>guide</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>g</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>H</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This loss encourages the model to produce sentiment representations that are semantically aligned with external guidance, thereby leveraging high-level knowledge to improve robustness and accuracy during multimodal fusion. This design bridges implicit commonsense reasoning from LLMs with low-level multimodal representations, enabling the model to handle sentiment ambiguity more effectively.</p>
<p>We employ a multi-task training strategy integrating image&#x2013;aspect relevance prediction, region-level alignment, sentiment classification, and external knowledge guidance to fully exploit multimodal signals at multiple granularities and enhance generalization via language model priors.</p>
<p>The overall training objective is formulated as a weighted sum of all sub-losses:<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mrow><mml:mrow><mml:mi>&#x1D4A5;</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>MASC</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>ATT</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>guide</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>MASC</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the primary loss for multimodal aspect-based sentiment classification, <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RE</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> supervises the prediction of image-aspect relevance, <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>ATT</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> enforces semantic alignment between image regions and aspect terms via region-level attention, <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mrow><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>guide</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> promotes consistency between the fused sentiment representation and external language model guidance.</p>
<p>The coefficients <italic>&#x03BB;</italic><sub>1</sub>, <italic>&#x03BB;</italic><sub>2</sub>, <italic>&#x03BB;</italic><sub>3</sub> are hyperparameters that control the relative contribution of each auxiliary loss. These values are selected based on validation performance to ensure that the auxiliary objectives enhance&#x2014;but do not overpower&#x2014;the primary sentiment classification task.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results and Analysis</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>In this study, experiments are conducted on the widely used Twitter-15 [<xref ref-type="bibr" rid="ref-43">43</xref>] and Twitter-17 [<xref ref-type="bibr" rid="ref-44">44</xref>] datasets, where each multimodal tweet comprises a textual post, a corresponding image, one or more aspect terms, and sentiment labels (negative, neutral, positive) for each aspect, as detailed in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Detailed information of Twitter-15 and Twitter-17 datasets</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Label</th>
<th colspan="3">Twitter-15</th>
<th colspan="3">Twitter-17</th>
</tr>
<tr>
<th align="center">Train</th>
<th align="center">Dev</th>
<th align="center">Test</th>
<th align="center">Train</th>
<th align="center">Dev</th>
<th align="center">Test</th>
</tr>
</thead>
<tbody>
<tr>
<td>Positive</td>
<td>928</td>
<td>303</td>
<td>317</td>
<td>1508</td>
<td>515</td>
<td align="center">493</td>
</tr>
<tr>
<td>Neutral</td>
<td>1883</td>
<td>670</td>
<td>607</td>
<td>1638</td>
<td>517</td>
<td align="center">573</td>
</tr>
<tr>
<td>Negative</td>
<td>368</td>
<td>149</td>
<td>113</td>
<td>416</td>
<td>144</td>
<td align="center">168</td>
</tr>
<tr>
<td>Total</td>
<td>3179</td>
<td>1122</td>
<td>1037</td>
<td>3562</td>
<td>1176</td>
<td align="center">1234</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Hyperparameter Settings and Evaluation Metrics</title>
<p>We use the pre-trained RoBERTa model with 24 layers, a hidden size of 1024, and 355 million parameters for textual and aspect representations, and Faster R-CNN to extract 100 region-of-interest features with a dimensionality of 2048 from images. Optimization uses BertAdam with learning rates of 1 &#x00D7; 10<sup>&#x2212;5</sup> for sentiment classification, 1 &#x00D7; 10<sup>&#x2212;6</sup> for visual region alignment, and 5 &#x00D7; 10<sup>&#x2212;6</sup> for external knowledge guidance, with a warmup proportion of 0.1. Training runs for 9 epochs with a batch size of 32 on a 32 GB NVIDIA V100 GPU. Performance is evaluated on the test set using Accuracy and Macro-F1 for baseline comparison.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Baseline Models</title>
<p>To assess the performance of aspect-based sentiment classification, we compare the proposed model with three categories of baseline models: image-only models, text-only models, and multimodal models.</p>
<p><bold>Image-only model:</bold> Res-Aspect [<xref ref-type="bibr" rid="ref-20">20</xref>] uses ResNet and BERT to extract visual and textual features.</p>
<p><bold>Text-only models:</bold> Representative examples include MGAN [<xref ref-type="bibr" rid="ref-19">19</xref>] with multi-granularity attention, BERT [<xref ref-type="bibr" rid="ref-4">4</xref>] and BART [<xref ref-type="bibr" rid="ref-5">5</xref>] as pre-trained language models for sentiment prediction, ATAE-LSTM [<xref ref-type="bibr" rid="ref-2">2</xref>] and IAN [<xref ref-type="bibr" rid="ref-3">3</xref>] with attention-based aspect modeling, RAM [<xref ref-type="bibr" rid="ref-17">17</xref>] using weighted memory, and TNet [<xref ref-type="bibr" rid="ref-18">18</xref>] applying CNN layers to extract salient features.</p>
<p><bold>Multimodal models:</bold> Multimodal models integrate textual and visual features through attention, memory, or knowledge-enhanced mechanisms for sentiment classification. Examples include cross-modal attention networks such as Res-MGAN and Res-BERT, memory-based architectures like MIMN [<xref ref-type="bibr" rid="ref-11">11</xref>], entity-sensitive fusion models such as ESAFN [<xref ref-type="bibr" rid="ref-13">13</xref>], and pre-trained vision&#x2013;language transformers like TomBERT [<xref ref-type="bibr" rid="ref-12">12</xref>] and ViLBERT [<xref ref-type="bibr" rid="ref-22">22</xref>]. Other approaches enhance visual grounding or sentiment cues via caption-based modeling in CapTrBERT [<xref ref-type="bibr" rid="ref-23">23</xref>], saliency detection in SaliencyBERT [<xref ref-type="bibr" rid="ref-22">22</xref>], semantic gap reduction in HIMT [<xref ref-type="bibr" rid="ref-14">14</xref>], knowledge integration in KEF [<xref ref-type="bibr" rid="ref-15">15</xref>], facial expression cues in FITE [<xref ref-type="bibr" rid="ref-25">25</xref>], emotional region detection in ARFN [<xref ref-type="bibr" rid="ref-26">26</xref>], and global&#x2013;local feature fusion with co-attention in GLFFCA [<xref ref-type="bibr" rid="ref-16">16</xref>]. More recent advances include MDCA [<xref ref-type="bibr" rid="ref-27">27</xref>], which formulates MABSA as dual-cause analysis guided by large language model reasoning, and DQPSA [<xref ref-type="bibr" rid="ref-28">28</xref>], which leverages dual-query prompting with an energy-based expert mechanism.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experimental Results</title>
<p>We conduct several comparative experiments based on the benchmark datasets, Twitter-15 and Twitter-17, to assess the effectiveness of the proposed Gated-Linear Aspect-Aware Multimodal Sentiment Network (GLAMSNet).</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Comparitive Analysis</title>
<p>The bold entries in <xref ref-type="table" rid="table-3">Table 3</xref> indicate the optimal results of the obtained evaluations. <xref ref-type="table" rid="table-3">Table 3</xref> provides a performance comparison for the proposed GLAMSNet with several representative baseline models on the Twitter-15 and Twitter-17 datasets, with evaluations based on Accuracy (Acc) and Macro-F1 Score (F1). The baselines encompass text-only, image-only, and multimodal approaches. To ensure equitable comparison, all models are assessed under identical experimental settings.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance comparison on the Twitter-15 and Twitter-17 datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="2"></th>
<th colspan="2">Twitter-15</th>
<th colspan="2">Twitter-17</th>
</tr>
<tr>
<th>Modality</th>
<th>Methods</th>
<th>Acc</th>
<th>F1</th>
<th>Acc</th>
<th>F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Image-only</td>
<td>Res-Aspect</td>
<td>59.49</td>
<td>47.79</td>
<td>57.86</td>
<td>53.98</td>
</tr>
<tr>
<td rowspan="3">Text-only</td>
<td>MGAN</td>
<td>71.20</td>
<td>64.20</td>
<td>64.80</td>
<td>61.50</td>
</tr>
<tr>
<td>BERT</td>
<td>74.30</td>
<td>70.00</td>
<td>68.90</td>
<td>66.10</td>
</tr>
<tr>
<td>BART</td>
<td>76.00</td>
<td>67.60</td>
<td>69.50</td>
<td>67.00</td>
</tr>
<tr>
<td rowspan="17">Multimodal (Text &#x002B; Image)</td>
<td>Res-MGAN</td>
<td>71.70</td>
<td>63.90</td>
<td>66.40</td>
<td>63.00</td>
</tr>
<tr>
<td>Res-Bert</td>
<td>75.02</td>
<td>69.21</td>
<td>69.20</td>
<td>66.48</td>
</tr>
<tr>
<td>MIMN</td>
<td>71.84</td>
<td>65.69</td>
<td>65.88</td>
<td>62.99</td>
</tr>
<tr>
<td>ESAFN</td>
<td>73.38</td>
<td>67.37</td>
<td>67.83</td>
<td>64.22</td>
</tr>
<tr>
<td>ViLBERT</td>
<td>73.76</td>
<td>69.85</td>
<td>67.42</td>
<td>64.87</td>
</tr>
<tr>
<td>TomBERT(resnet)</td>
<td>76.60</td>
<td>71.57</td>
<td>69.42</td>
<td>67.70</td>
</tr>
<tr>
<td>TomBERT(rcnn)</td>
<td>77.03</td>
<td>72.85</td>
<td>69.77</td>
<td>67.59</td>
</tr>
<tr>
<td>CapTrBERT</td>
<td>78.01</td>
<td>73.25</td>
<td>69.77</td>
<td>68.42</td>
</tr>
<tr>
<td>saliencyBERT</td>
<td>77.03</td>
<td>72.36</td>
<td>69.69</td>
<td>67.59</td>
</tr>
<tr>
<td>HIMT</td>
<td>78.41</td>
<td>73.68</td>
<td>71.14</td>
<td>69.16</td>
</tr>
<tr>
<td>KEF-saliencyBERT</td>
<td>78.15</td>
<td>73.54</td>
<td>71.88</td>
<td>68.94</td>
</tr>
<tr>
<td>FITE</td>
<td>78.49</td>
<td>73.90</td>
<td>70.09</td>
<td>68.70</td>
</tr>
<tr>
<td>ARFN</td>
<td>78.50</td>
<td>73.70</td>
<td>70.58</td>
<td>68.43</td>
</tr>
<tr>
<td>GLFFCA</td>
<td>77.72</td>
<td>74.21</td>
<td>71.15</td>
<td>69.45</td>
</tr>
<tr>
<td>MDCA</td>
<td>78.74</td>
<td>74.23</td>
<td>73.51</td>
<td>71.25</td>
</tr>
<tr>
<td>DQPSA</td>
<td>79.02</td>
<td>74.86</td>
<td>73.94</td>
<td>71.82</td>
</tr>
<tr>
<td><bold>Our model</bold></td>
<td><bold>79.36</bold></td>
<td><bold>74.72</bold></td>
<td><bold>74.31</bold></td>
<td><bold>72.01</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results show that GLAMSNet consistently outperforms most existing methods on both datasets. Compared with the strongest baseline group (GLFFCA, MDCA, and DQPSA) [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>], our model yields further improvements. Specifically, GLAMSNet achieves F1 gains of 0.51% on Twitter-15 and 2.56% on Twitter-17 compared with GLFFCA [<xref ref-type="bibr" rid="ref-16">16</xref>], while also surpassing the more recent MDCA [<xref ref-type="bibr" rid="ref-27">27</xref>]and DQPSA [<xref ref-type="bibr" rid="ref-28">28</xref>]. MDCA frames MABSA as dual-cause analysis with large language model prompting, reaching 78.74% accuracy and 74.23% F1 on Twitter-15, and 73.51% accuracy and 71.25% F1 on Twitter-17. DQPSA introduces a dual-query prompt with an energy-based expert mechanism and delivers 79.02% accuracy and 74.86% F1 on Twitter-15, and 73.94% accuracy and 71.82% F1 on Twitter-17. Despite their competitiveness, GLAMSNet consistently surpasses both, achieving margins of 0.34%&#x2013;0.82% in accuracy and 0.28%&#x2013;1.20% in F1. This confirms the effectiveness of combining hierarchical alignment supervision, gated-linear aspect-aware fusion, multi-head linear attention, and external knowledge guidance into a unified framework.</p>
<p>Finally, EMT-DLFR [<xref ref-type="bibr" rid="ref-29">29</xref>] introduces efficient multimodal transformers with dual-level feature restoration for video-based sentiment analysis. As it addresses different modalities and datasets, direct comparison is infeasible. Nevertheless, its global&#x2013;local fusion and restoration strategies provide promising directions for extending GLAMSNet to video sentiment analysis in future work.</p>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Model Performance Analysis</title>
<p>To evaluate GLAMSNet&#x2019;s performance, we track Accuracy and Macro-F1 across epochs 0&#x2013;8 on the validation and test sets of Twitter-15 and Twitter-17, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. On validation sets, accuracy rises rapidly in early epochs&#x2014;63.73% to 74.96% on Twitter-15 and 43.88% to 73.13% on Twitter-17&#x2014;before stabilizing, with similar trends in F1 scores. On test sets, accuracy peaks at 79.56% for Twitter-15 and 74.31% for Twitter-17, while F1 reaches 75.14% and 73.02%, respectively, maintaining stable performance thereafter. GLAMSNet converges slightly faster on Twitter-15, but its peak results on Twitter-17 demonstrate adaptability to complex data. These trends confirm the effectiveness of the GLAF module and external guidance in achieving robust, high-accuracy sentiment prediction.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Performance of GLAMSNet on Twitter-15 and Twitter-17 across training epochs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_71656-fig-3.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Ablation Study</title>
<p>To validate the contribution of each component in GLAMSNet, we perform ablation studies on Twitter-15 and Twitter-17 by removing Knowledge-Guided, MHLAttention, GatedFusion, and GLAF modules individually. As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, &#x201C;w/o Knowledge-Guided&#x201D; excludes external knowledge guidance, &#x201C;w/o MHLAttention&#x201D; replaces it with standard Transformer self-attention, and &#x201C;w/o GatedFusion&#x201D; adopts the baseline gate-based fusion from the first-stage architecture without our enhanced design, while the &#x201C;Full Model&#x201D; denotes the complete architecture. In addition, we include fine-grained variants to analyze subcomponents: &#x201C;w/o ReLU in MHLAttention&#x201D; and &#x201C;w/o Linear in MHLAttention&#x201D; ablate the activation and projection operations within the linear attention module, respectively; &#x201C;w/o Contextual Prompting&#x201D; removes the contextualized prompt design for external knowledge guidance; and &#x201C;w/o GPT-4o Guidance&#x201D; substitutes GPT-4o&#x2013;derived priors with outputs from GPT-2.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation study results on the Twitter-15 and Twitter-17 datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Modality</th>
<th>Tw15-Acc</th>
<th>Tw15-F1</th>
<th>Tw17-Acc</th>
<th>Tw17-F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>Full Model</td>
<td> <bold>79.36</bold></td>
<td><bold>74.72</bold></td>
<td><bold>74.31</bold></td>
<td><bold>72.01</bold></td>
</tr>
<tr>
<td>w/o Knowledge-Guided</td>
<td>77.65</td>
<td>72.58</td>
<td>72.45</td>
<td>70.35</td>
</tr>
<tr>
<td>w/o GPT-4o Guidance</td>
<td>78.10</td>
<td>73.05</td>
<td>72.89</td>
<td>70.71</td>
</tr>
<tr>
<td>w/o Contextual Prompting</td>
<td>78.24</td>
<td>73.21</td>
<td>73.02</td>
<td>70.92</td>
</tr>
<tr>
<td>w/o MHLAttention</td>
<td>78.46</td>
<td>73.85</td>
<td>73.52</td>
<td>71.22</td>
</tr>
<tr>
<td>w/o ReLU in MHLAttention</td>
<td>78.61</td>
<td>73.87</td>
<td>73.36</td>
<td>71.18</td>
</tr>
<tr>
<td>w/o Linear in MHLAttention</td>
<td>78.49</td>
<td>73.79</td>
<td>73.25</td>
<td>71.09</td>
</tr>
<tr>
<td>w/o GatedFusion</td>
<td>78.76</td>
<td>74.05</td>
<td>73.85</td>
<td>71.55</td>
</tr>
<tr>
<td>w/o GLAF</td>
<td>77.62</td>
<td>73.32</td>
<td>72.25</td>
<td>70.01</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The ablation studies demonstrate the effectiveness of key components. First, removing the Knowledge-Guided module reduces accuracy by 1.71 points and F1-score by 2.14 points on Twitter-15, and 1.86 and 1.66 points on Twitter-17, confirming the value of GPT-predicted sentiment priors. In addition, replacing GPT-4o with GPT-2 as the external knowledge source reduces accuracy by 1.26 points and F1-score by 1.43 points on Twitter-15, while removing contextual prompting decreases accuracy by 1.12 points and F1-score by 1.09 points on Twitter-17, further demonstrating the necessity of both high-quality external knowledge and effective prompt design. Second, replacing the Multi-Head Linear Attention module decreases performance by 0.79 to 0.90 points, and removing either the ReLU or Linear transformation within this module causes additional drops of about 0.9 to 1.1 points, underscoring the importance of both activation and projection operations. Third, removing the Gated Fusion module leads to drops of 0.46 to 0.67 points, reflecting its role in balancing cross-modal information. Most significantly, ablating the Gated-Linear Aspect-Aware Fusion module causes the largest declines (2.00 to 2.40 points), underscoring its importance in mitigating aspect-image misalignment, particularly given that 58% of dataset images are aspect-irrelevant.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Visual Region Analysis</title>
<p>To evaluate GLAMSNet&#x2019;s region selection and prediction in MASC, we examine three samples from Twitter-15 and Twitter-17 covering athletic achievements, sports conflicts, and entertainment performances, representing different cross-modal alignment levels.</p>
<sec id="s4_6_1">
<title>Illustration Analysis</title>
<p>As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> (3 &#x00D7; 3 grid layout), each sample presents the original image, attention-weighted regions, and predicted (red) vs. ground-truth (blue) bounding boxes. The first row in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows a sports achievement from Twitter-15, depicting Kevin Durant and Russell Westbrook in a high-scoring game. The attention map [<xref ref-type="bibr" rid="ref-45">45</xref>] focuses on faces and hand gestures but partially covers the background, resulting in an IoU of 0.617, indicating background interference in dynamic scenes. The second row shows a sports honor from Twitter-17, where Kevin Durant receives the NBA Championship and Finals MVP awards. The model attends to the player, trophy, and podium with minimal background noise, achieving an IoU of 0.722 and reliable spatial focus in positive scenarios. The third row presents an entertainment performance from Twitter-17, featuring Mark Ronson and Lady Gaga at the Met Gala. The model focuses on both performers and the central stage with little background attention, yielding an IoU of 0.895, which reflects precise spatial focus and strong cross-modal alignment for accurate sentiment localization.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Visual region analysis of GLAMSNet on Twitter-15 and Twitter-17</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_71656-fig-4.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Case Studies</title>
<p>Three case studies compare our model with TomBERT [<xref ref-type="bibr" rid="ref-12">12</xref>] and GLFFCA [<xref ref-type="bibr" rid="ref-16">16</xref>] on textual&#x2013;visual samples with multiple aspect sentiments, covering sports memorials, conflicts, and celebrations. These examples highlight the benefits of multimodal fusion and the limitations of baselines, as summarized in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Case studies on Twitter-15 and Twitter-17</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Text</th>
<th align="center">Three years ago today&#x2014;<sup>1</sup>Robin van Persie<sub>neu</sub>: The Flying Dutchman.</th>
<th align="center"><sup>1</sup>Kevin Durant<sub>neg</sub> and <sup>2</sup>LeBron James<sub>neg</sub> had some tense words after Kevin Love&#x2019;s foul</th>
<th align="center"><sup>1</sup>Lewis Hamilton<sub>neu</sub> celebrates with <sup>2</sup>Justin Bieber<sub>pos</sub> after winning thrilling Monaco Grand Prix race</th>
</tr>
</thead>
<tbody>
<tr>
<td>Image</td>
<td><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMC_71656-inline-1.tif"/></td>
<td><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMC_71656-inline-2.tif"/></td>
<td><inline-graphic mimetype="image" mime-subtype="png" xlink:href="CMC_71656-inline-3.tif"/></td>
</tr>
<tr>
<td rowspan="2">TomBERT</td>
<td rowspan="2">Robin van Persie-pos</td>
<td>Kevin Durant-neg</td>
<td>Lewis Hamilton-neg</td>
</tr>
<tr>
<td>LeBron James-neg</td>
<td>Justin Bieber-neu</td>
</tr>
<tr>
<td rowspan="2">GLFFCA</td>
<td rowspan="2">Robin van Persie-pos</td>
<td>Kevin Durant-neg</td>
<td>Lewis Hamilton-neg</td>
</tr>
<tr>
<td>LeBron James-neg</td>
<td>Justin Bieber-neu</td>
</tr>
<tr>
<td rowspan="2"><bold>Ours</bold></td>
<td rowspan="2">Robin van Persie-neu</td>
<td>KevinDurant-neg</td>
<td>Lewis Hamilton-neu</td>
</tr>
<tr>
<td>LeBron James-neg</td>
<td>Justin Bieber-pos</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-5">Table 5</xref>, three case studies compare our model with TomBERT [<xref ref-type="bibr" rid="ref-12">12</xref>] and GLFFCA [<xref ref-type="bibr" rid="ref-16">16</xref>]. In the sports memorial scenario, Robin van Persie&#x2019;s &#x201C;Flying Dutchman&#x201D; goal is neutral but misclassified as positive by TomBERT and GLFFCA, whereas our model uses multimodal context to resolve the ambiguity. In the sports conflict scenario, a foul-induced confrontation between Kevin Durant and LeBron James is correctly predicted as negative by all models, with ours showing higher stability. In the event celebration scenario, Lewis Hamilton is neutral and Justin Bieber is positive; only our model correctly distinguishes both via fine-grained multimodal fusion. These cases confirm our approach&#x2019;s effectiveness, with future work aimed at refining cross-modal attention for better generalization.</p>

</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, we propose a unified framework for multimodal aspect-based sentiment analysis, named Gated-Linear Aspect-Aware Multimodal Sentiment Network (GLAMSNet), which is designed to address three key challenges in MABSA: insufficient aspect&#x2013;image alignment, modality imbalance, and limited semantic reasoning. To this end, the model integrates hierarchical alignment supervision, aspect-aware gated fusion, and external language model guidance into a cohesive architecture. These components work collaboratively to enhance cross-modal interaction, improve representation quality, and enable more accurate and robust sentiment inference. Experimental results on Twitter-15 and Twitter-17 demonstrate the effectiveness of our approach, achieving 79.36% accuracy and 74.72% F1-score, and 74.31% accuracy and 72.01% F1-score, respectively&#x2014;outperforming a range of strong baselines.</p>
<p>Despite its effectiveness, GLAMSNet relies on manual alignment annotations, limiting scalability in low-resource settings, and may be affected by noisy external knowledge that misleads sentiment inference. In future studies, we are interested in exploring self-supervised or weakly supervised alignment strategies, developing adaptive mechanisms to assess and calibrate the reliability of external knowledge sources, and further investigating the quality and controllability of knowledge generated by large language models.</p>
</sec>
</body>
<back>
<ack>
<p>The authors are deeply grateful to all team members involved in this research.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported in part by the National Nature Science Foundation of China under Grants 62476216 and 62273272, in part by the Key Research and Development Program of Shaanxi Province under Grant 2024GX-YBXM-146, in part by the Scientific Research Program Funded by Education Department of Shaanxi Provincial Government under Grant 23JP091, and the Youth Innovation Team of Shaanxi Universities.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization: Dan Wang and Zhoubin Li; Methodology: Dan Wang, Zhoubin Li and Yuze Xia; Formal analysis and investigation: Dan Wang, Zhoubin Li and Yuze Xia; Writing&#x2014;original draft preparation: Dan Wang and Zhoubin Li; Writing&#x2014;review and editing: Dan Wang and Zhoubin Li; Funding acquisition: Dan Wang and Zhenhua Yu; Supervision: Zhenhua Yu. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on reasonable request from the authors.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nazir</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Issues and challenges of aspect-based sentiment analysis: a comprehensive survey</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2022</year>;<volume>13</volume>(<issue>2</issue>):<fpage>845</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAFFC.2020.2970399</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Attention-based LSTM for aspect-level sentiment classification</article-title>. In: <conf-name>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing; 2016 Nov 1&#x2013;5</conf-name>; <publisher-loc>Austin, TX, USA</publisher-loc>. p. <fpage>606</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/d16-1058</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Interactive attention networks for aspect-level sentiment classification</article-title>. In: <conf-name>Proceedings of the 26th International Joint Conference on Artificial Intelligence; 2017 Aug 19&#x2013;26</conf-name>; <publisher-loc>Melbourne, VIC, Australia</publisher-loc>. p. <fpage>4068</fpage>&#x2013;<lpage>74</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2017/568</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name></person-group>. <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. <comment>arXiv:1810.04805. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1810.04805</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lewis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ghazvininejad</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Levy</surname> <given-names>O</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension</article-title>. <comment>arXiv:1910.13461. 2019</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1910.13461</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Low-rank adapter layers and bidirectional gated feature fusion for multimodal hateful memes classification</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>84</volume>(<issue>1</issue>):<fpage>1863</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.064734</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Song</surname> <given-names>K</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fake news detection based on cross-modal ambiguity computation and multi-scale feature fusion</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>83</volume>(<issue>2</issue>):<fpage>2659</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.060025</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>LA</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Multimodal aspect-based sentiment analysis: a survey of tasks, methods, challenges and future directions</article-title>. <source>Inf Fusion</source>. <year>2024</year>;<volume>112</volume>(<issue>3</issue>):<fpage>102552</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2024.102552</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rahman</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mazrur</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mia</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Idress Ekfa</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Exploring the effectiveness of machine learning and deep learning algorithms for sentiment analysis: a systematic literature review</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>84</volume>(<issue>3</issue>):<fpage>4105</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.066910</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Javed</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shoaib</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jaleel</surname> <given-names>A</given-names></string-name>, <string-name><surname>Deriche</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nawaz</surname> <given-names>S</given-names></string-name></person-group>. <article-title>X-OODM: leveraging explainable object-oriented design methodology for multi-domain sentiment analysis</article-title>. <source>Comput Mater Contin</source>. <year>2025</year>;<volume>82</volume>(<issue>3</issue>):<fpage>4977</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2025.057359</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Multi-interactive memory network for aspect based multimodal sentiment analysis</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2019</year>;<volume>33</volume>(<issue>1</issue>):<fpage>371</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v33i01.3301371</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Adapting BERT for target-oriented multimodal sentiment classification</article-title>. In: <conf-name>Proceedings of the 28th International Joint Conference on Artificial Intelligence; 2019 Aug 10&#x2013;16</conf-name>; <publisher-loc>Macao, China</publisher-loc>. p. <fpage>5408</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2019/751</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Entity-sensitive attention and fusion network for entity-level multimodal sentiment classification</article-title>. <source>IEEE/ACM Trans Audio Speech Lang Process</source>. <year>2020</year>;<volume>28</volume>:<fpage>429</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TASLP.2019.2957872</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Hierarchical interactive multimodal transformer for aspect-based multimodal sentiment analysis</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2023</year>;<volume>14</volume>(<issue>3</issue>):<fpage>1966</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAFFC.2022.3171091</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Long</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Learning from adjective-noun pairs: a knowledge-enhanced framework for target-oriented multimodal sentiment classification</article-title>. In: <conf-name>Proceedings of the 29th International Conference on Computational Linguistics; 2022 Oct 12&#x2013;17</conf-name>; <publisher-loc>Gyeongju, Republic of Korea</publisher-loc>. p. <fpage>6784</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Aspect-level multimodal sentiment analysis based on co-attention fusion</article-title>. <source>Int J Data Sci Anal</source>. <year>2025</year>;<volume>20</volume>(<issue>2</issue>):<fpage>903</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s41060-023-00497-3</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bing</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Recurrent attention network on memory for aspect sentiment analysis</article-title>. In: <conf-name>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing; 2017 Sep 7&#x2013;11</conf-name>; <publisher-loc>Copenhagen, Denmark</publisher-loc>. p. <fpage>452</fpage>&#x2013;<lpage>61</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/d17-1047</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Bing</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lam</surname> <given-names>W</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Transformation networks for target-oriented sentiment classification</article-title>. In: <conf-name>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2018 Jul 15&#x2013;20</conf-name>; <publisher-loc>Melbourne, Australia</publisher-loc>. p. <fpage>946</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/p18-1087</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>F</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Multi-grained attention network for aspect-level sentiment classification</article-title>. In: <conf-name>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing; 2018 Oct 31&#x2013;Nov 4</conf-name>; <publisher-loc>Brussels, Belgium</publisher-loc>. p. <fpage>3433</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/d18-1380</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Multi-level textual-visual alignment and fusion network for multimodal aspect-based sentiment analysis</article-title>. <source>Artif Intell Rev</source>. <year>2024</year>;<volume>57</volume>(<issue>4</issue>):<fpage>78</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10462-023-10685-z</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Batra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Parikh</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>S</given-names></string-name></person-group>. <article-title>ViLBERT: pretraining task-agnostic visiolinguistic representations for vision-and-language tasks</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2019</year>;<volume>32</volume>:<fpage>13</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Exploiting BERT for multimodal target sentiment classification through input space translation</article-title>. In: <conf-name>Proceedings of the 29th ACM International Conference on Multimedia; 2021 Oct 20&#x2013;24</conf-name>; <publisher-loc>Virtual</publisher-loc>. p. <fpage>3034</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3474085.3475692</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sheng</surname> <given-names>V</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>SaliencyBERT: recurrent attention network for target-oriented multimodal sentiment classification</article-title>. In: <conf-name>Chinese Conference on Pattern Recognition and Computer Vision (PRCV); 2021 Oct 29&#x2013;Nov 1</conf-name>; <publisher-loc>Beijing, China</publisher-loc>. p. <fpage>3</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-88010-1_1</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Face-sensitive image-to-emotional-text cross-modal translation for multimodal aspect-based sentiment analysis</article-title>. In: <conf-name>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing; 2022 Dec 7&#x2013;11</conf-name>; <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>. p. <fpage>3324</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2022.emnlp-main.219</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>H</given-names></string-name>, <string-name><surname>Al-Nabhan</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Affective region recognition and fusion network for target-level multimodal sentiment classification</article-title>. <source>IEEE Trans Emerg Top Comput</source>. <year>2024</year>;<volume>12</volume>(<issue>3</issue>):<fpage>688</fpage>&#x2013;<lpage>99</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TETC.2022.3231746</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Dual causes generation assisted model for multimodal aspect-based sentiment classification</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2025</year>;<volume>36</volume>(<issue>5</issue>):<fpage>9298</fpage>&#x2013;<lpage>312</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2024.3415028</pub-id>; <pub-id pub-id-type="pmid">38917280</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A novel energy based model mechanism for multi-modal aspect-based sentiment analysis</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2024</year>;<volume>38</volume>(<issue>17</issue>):<fpage>18869</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i17.29852</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Efficient multimodal transformer with dual-level feature restoration for robust multimodal sentiment analysis</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>309</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAFFC.2023.3274829</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lian</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Schuller</surname> <given-names>BW</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Contrastive learning based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities</article-title>. <source>IEEE Trans Affect Comput</source>. <year>2024</year>;<volume>15</volume>(<issue>4</issue>):<fpage>1856</fpage>&#x2013;<lpage>73</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAFFC.2024.3378570</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>LH</given-names></string-name></person-group>. <article-title>Exploit the visual sentiment of the item images to fuse with textual sentiment in context aware collaborative filtering</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>265</volume>(<issue>6</issue>):<fpage>125970</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.125970</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ott</surname> <given-names>M</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Du</surname> <given-names>J</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>RoBERTa: a robustly optimized BERT pretraining approach</article-title>. <comment>arXiv:1907.11692. 2019</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1907.11692</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Faster R-CNN: towards real-time object detection with region proposal networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2017</year>;<volume>39</volume>(<issue>6</issue>):<fpage>1137</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id>; <pub-id pub-id-type="pmid">27295650</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name></person-group>. <article-title>DeBERTa: decoding-enhanced BERT with disentangled attention</article-title>. <comment>arXiv:2006.03654. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2006.03654</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Beyer</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kolesnikov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Weissenborn</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>An image is worth 16 &#x00D7; 16 words: transformers for image recognition at scale</article-title>. <comment>arXiv:2010.11929. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2010.11929</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning transferable visual models from natural language supervision</article-title>. <comment>arXiv:2103.00020. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2103.00020</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Blip: bootstrapping language-image pre-training for unified vision-language understanding and generation</article-title>. In: <conf-name>International Conference on Machine learning; 2022 Jul 17&#x2013;23</conf-name>; <publisher-loc>Baltimore, MD, USA</publisher-loc>. p. <fpage>12888</fpage>&#x2013;<lpage>900</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Targeted multimodal sentiment classification based on coarse-to-fine grained image-target matching</article-title>. In: <conf-name>Proceedings of the 31st International Joint Conference on Artificial Intelligence; 2022 Jul 23&#x2013;29</conf-name>; <publisher-loc>Vienna, Austria</publisher-loc>. p. <fpage>4482</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2022/622</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tsai</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>PP</given-names></string-name>, <string-name><surname>Kolter</surname> <given-names>JZ</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>LP</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Multimodal transformer for unaligned multimodal language sequences</article-title>. In: <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; 2019 Jul 28&#x2013;Aug 2</conf-name>; <publisher-loc>Florence, Italy</publisher-loc>. p. <fpage>6559</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/p19-1656</pub-id>; <pub-id pub-id-type="pmid">32362720</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Rethinking diversified and discriminative proposal generation for visual grounding</article-title>. In: <conf-name>Proceedings of the 27th International Joint Conference on Artificial Intelligence; 2018 Jul 13&#x2013;19</conf-name>; <publisher-loc>Stockholm, Sweden</publisher-loc>. p. <fpage>1114</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2018/155</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Berg</surname> <given-names>T</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>M</given-names></string-name></person-group>. <article-title>TVQA&#x002B;: spatio-temporal grounding for video question answering</article-title>. In: <conf-name> Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; 2020 Jul 5&#x2013;10</conf-name>; <publisher-loc>Online</publisher-loc>. p. <fpage>8211</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2020.acl-main.730</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name></person-group>. <article-title>EfficientViT: lightweight multi-scale attention for high-resolution dense prediction</article-title>. In: <conf-name>2023 IEEE/CVF International Conference on Computer Vision (ICCV); 2023 Oct 1&#x2013;6</conf-name>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>17256</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv51070.2023.01587</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Neves</surname> <given-names>L</given-names></string-name>, <string-name><surname>Carvalho</surname> <given-names>V</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Visual attention model for name tagging in multimodal social media</article-title>. In: <conf-name>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2018 Jul 15&#x2013;20</conf-name>; <publisher-loc>Melbourne, Australia</publisher-loc>. p. <fpage>1990</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/p18-1185</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Adaptive co-attention network for named entity recognition in tweets</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2018</year>;<volume>32</volume>(<issue>1</issue>):<fpage>5674</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v32i1.11962</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Leem</surname> <given-names>S</given-names></string-name>, <string-name><surname>Seo</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Attention guided CAM: visual explanations of vision transformer guided by self-attention</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2024</year>;<volume>38</volume>(<issue>4</issue>):<fpage>2956</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i4.28077</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>