<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">59745</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.059745</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>UniTrans: Unified Parameter-Efficient Transfer Learning and Multimodal Alignment for Large Multimodal Foundation Model</article-title>
<alt-title alt-title-type="left-running-head">UniTrans: Unified Parameter-Efficient Transfer Learning and Multimodal Alignment for Large Multimodal Foundation Model</alt-title>
<alt-title alt-title-type="right-running-head">UniTrans: Unified Parameter-Efficient Transfer Learning and Multimodal Alignment for Large Multimodal Foundation Model</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Sun</surname><given-names>Jiakang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Ke</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>He</surname><given-names>Xinyang</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Xu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Ke</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Peng</surname><given-names>Cheng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>pengchengcasit@163.com</email></contrib>
<aff id="aff-1"><label>1</label><institution>Chengdu Institute of Computer Application, Chinese Academy of Sciences</institution>, <addr-line>Chengdu, 610213</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computer Science and Technology, University of Chinese Academy of Sciences</institution>, <addr-line>Beijing, 101499</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Cheng Peng. Email: <email>pengchengcasit@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>26</day><month>03</month><year>2025</year>
</pub-date>
<volume>83</volume>
<issue>1</issue>
<fpage>219</fpage>
<lpage>238</lpage>
<history>
<date date-type="received">
<day>16</day>
<month>10</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>1</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_59745.pdf"></self-uri>
<abstract>
<p>With the advancements in parameter-efficient transfer learning techniques, it has become feasible to leverage large pre-trained language models for downstream tasks under low-cost and low-resource conditions. However, applying this technique to multimodal knowledge transfer introduces a significant challenge: ensuring alignment across modalities while minimizing the number of additional parameters required for downstream task adaptation. This paper introduces UniTrans, a framework aimed at facilitating efficient knowledge transfer across multiple modalities. UniTrans leverages Vector-based Cross-modal Random Matrix Adaptation to enable fine-tuning with minimal parameter overhead. To further enhance modality alignment, we introduce two key components: the Multimodal Consistency Alignment Module and the Query-Augmentation Side Network, specifically optimized for scenarios with extremely limited trainable parameters. Extensive evaluations on various cross-modal downstream tasks demonstrate that our approach surpasses state-of-the-art methods while using just 5% of their trainable parameters. Additionally, it achieves superior performance compared to fully fine-tuned models on certain benchmarks.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Parameter-efficient transfer learning</kwd>
<kwd>multimodal alignment</kwd>
<kwd>image captioning</kwd>
<kwd>image-text retrieval</kwd>
<kwd>visual question answering</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The current paradigm in artificial intelligence has shifted from developing domain-specific models to pretraining large models on extensive datasets, followed by fine-tuning for downstream tasks [<xref ref-type="bibr" rid="ref-1">1</xref>]. This shift has led to the development of several prominent large pre-trained models, such as LLaMA [<xref ref-type="bibr" rid="ref-2">2</xref>], SAM [<xref ref-type="bibr" rid="ref-3">3</xref>] and BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]. However, as the number of parameters in foundational pre-trained models continues to grow (such as the 175B parameters in GPT-3 [<xref ref-type="bibr" rid="ref-5">5</xref>]), the computational and storage resources required for full-parameter fine-tuning have increased significantly. Consequently, it has become crucial to identify methods that strike an effective balance between cost efficiency and fine-tuning performance.</p>
<p>Transfer learning has effectively addressed the challenge of applying knowledge from one task to another related task, enhancing learning efficiency and generalization ability [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>]. In particular, parameter-efficient transfer learning has attracted particular attention, as it enables knowledge transfer by adding only a small number of additional training parameters, thereby drastically reducing computational and storage requirements [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Current mainstream efficient-parameter transfer methods can be primarily categorized into prompt tuning, adapter tuning, and selective tuning. These methods either freeze the backbone and fine-tune only the added extra parameters or select a small subset of parameters from the backbone for training. This approach significantly reduces computational costs while achieving performance comparable to full fine-tuning. This method has demonstrated substantial success in both natural language processing and computer vision. Recently, there has been growing interest in multimodal foundational models. However, many of these studies either target only a single downstream task, neglecting the diverse range of tasks in multimodal models [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>], or they directly add learnable parameters using unimodal models&#x2019; parameter-efficient fine-tuning methods ignoring the alignment between different modalities [<xref ref-type="bibr" rid="ref-13">13</xref>]. Alternatively, some introduce excessive redundant parameters [<xref ref-type="bibr" rid="ref-14">14</xref>]. These approaches fail to solve the core problem in parameter-efficient fine-tuning for multimodal models: <bold><italic>achieving downstream knowledge transfer with minimal additional parameters while ensuring proper alignment between modalities</italic>.</bold></p>
<p>To address this issue, we design a novel and effective framework, UniTrans, to facilitate cross-modal knowledge transfer. Although efficient parameter transfer learning techniques, which involve adding extra parameters, have been widely applied in natural language processing and computer vision, there is still a lack of sufficient exploration in the multimodal domain. Based on extensive experiments, we propose <bold>vector-based cross-modal random matrix adaptation (VCRA)</bold>, which leverages Low-Rank Adaptation (LoRA) to decompose low-rank matrices into learnable scaling vectors and shared low-rank random matrices. VCRA employs a pair of random matrices for weight sharing, allowing fine-grained information between the image and text modalities to interact, thereby enhancing the visual-language modality representation.</p>
<p>Furthermore, during the fine-tuning process of multimodal base models on downstream tasks, the originally aligned image and text features may become disrupted, causing them to shift within their respective feature domains. To address this, we design the <bold>multimodal consistency alignment module (MCAM)</bold> and the <bold>query-augmentation side network (QASN)</bold>, which serve as regularizers for feature alignment during the fine-tuning process. MCAM, from a contrastive learning perspective, constrains the similarity ranking consistency between image-text pairs by designing a simple yet effective loss function without introducing additional parameters. At the same time, we observe that when the fusion network of a multimodal model has too many layers, query information loss occurs, which impacts the fusion and alignment between modalities. To resolve this, we propose the lightweight QASN, which adaptively supplements query information at various layers of the fusion network through a side-network approach, preventing matching errors between images and text caused by information loss. Finally, we evaluate our method on multiple cross-modal benchmarks, and the results show that our approach requires only 5% of the training parameters compared to state-of-the-art methods, significantly reducing training costs while outperforming traditional methods. Additionally, when compared to full-parameter training methods, our approach achieves comparable or even better performance (<xref ref-type="fig" rid="fig-1">Fig. 1</xref>).</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Performance comparison between image-text retrieval (left) and VQA (right) tasks. The size of the bubble represents the number of trainable parameters. Our UniTrans has achieved competitive performance in both tasks</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-1.tif"/>
</fig>
<p>In summary, our contributions can be summarized as follows:
<list list-type="simple">
<list-item><label>1)</label><p>We propose a lightweight and effective framework, UniTrans, for efficient cross-modal parameter knowledge transfer. The framework consists of VCRA and two modules, MCAM and QASN, designed to enhance modality alignment during the fine-tuning process.</p></list-item>
<list-item><label>2)</label><p>Based on Low-Rank Adaptation, we design a more suitable adapter for multimodal models, VCRA, which facilitates modality interaction and modality-specific adaptation through shared random matrices and learnable scaling vectors.</p></list-item>
<list-item><label>3)</label><p>We design two modules, MCAM and QASN, to constrain modality alignment, further improving the performance of multimodal downstream tasks.</p></list-item>
<list-item><label>4)</label><p>We conduct experiments on multiple multimodal benchmarks. The results show that our method reduces the trainable parameters to 5% compared to state-of-the-art method without sacrificing performance, and it significantly outperforms traditional methods. These benchmark test results are of significant importance for future research.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Vision-Language Models</title>
<p>In recent years, increasing efforts have focused on applying Vision-Language Models (VLMs) pre-trained with large-scale image-text pairs to downstream tasks [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Unlike pre-trained large language models [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>], VLMs typically extract multimodal features through separate text and image encoders, and then align these features using fusion mechanisms such as contrastive learning [<xref ref-type="bibr" rid="ref-15">15</xref>], transformer modules [<xref ref-type="bibr" rid="ref-4">4</xref>], Q-Former [<xref ref-type="bibr" rid="ref-19">19</xref>], or MLP [<xref ref-type="bibr" rid="ref-20">20</xref>]. The fused features are applied to multimodal downstream tasks. BLIP integrates both an encoder and a decoder, enabling support for both multimodal alignment and multimodal generation tasks within a single foundational model. This paper explores efficient parameter transfer methods tailored for multimodal models based on the BLIP framework.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Parameter-Efficient Transfer Learning</title>
<p>Cross-modal alignment refers to establishing correspondences between information from different modalities, enabling machines to recognize and understand the same or related information across various modalities. Current cross-modal alignment methods can be broadly categorized into attention-based alignment [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>], large cross-modal model-based alignment [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>], parameter-free interaction-based alignment [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>], and structure-based alignment [<xref ref-type="bibr" rid="ref-29">29</xref>]. However, these methods are typically applied during the model training process. As the number of model parameters grows, the pre-training and fine-tuning paradigm has become predominant, highlighting the need for inter-modal alignment mechanisms specifically designed for the fine-tuning process. Based on this, we have designed Query-Augmentation Side Network and Multimodal Consistency Alignment Module, which can serve as regularizers for fine-grained cross-modal alignment during the fine-tuning process.</p>
<p>As the number of parameters in foundational pre-trained models continues to grow, the cost of full-parameter fine-tuning for downstream tasks has become increasingly prohibitive, drawing more attention from the engineering community to parameter-efficient transfer learning. This approach facilitates knowledge transfer in downstream tasks by adding a small number of additional parameters and can be broadly categorized into the following types: prompt tuning [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>], adapter tuning [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>], and selective tuning [<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. While these methods have achieved significant success in the field of natural language processing (NLP), they remain underexplored in the multimodal domain. Existing work either applies these methods to multimodal models without accounting for the alignment between different modalities or focuses solely on a single downstream task. A recent study, UniAdapter [<xref ref-type="bibr" rid="ref-14">14</xref>], pioneered parameter-efficient transfer learning for multimodal models but introduced excessive redundant parameters. In contrast, our proposed method, UniTrans, reduces the number of parameters by an order of magnitude, while delivering comparable or even superior performance across multiple downstream tasks.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Low-Rank Adaptation (LoRA)</title>
<p>Low-Rank Adaptation is a type of parameter-efficient transfer learning that involves adapter fine-tuning. Unlike methods that add adapters, LoRA approximates weight changes through low-rank matrices during fine-tuning, allowing it to merge seamlessly with pre-trained weights during inference without introducing extra computational overhead. This significantly reduces the computational and storage resources required for tuning, providing an innovative solution for large pre-trained models. Base on this, AdaLoRA [<xref ref-type="bibr" rid="ref-36">36</xref>] uses Singular Value Decomposition (SVD) decomposition during fine-tuning to selectively remove insignificant singular values, dynamically adjusting the rank of the low-rank matrix for more efficient updates. Tied-LoRA [<xref ref-type="bibr" rid="ref-37">37</xref>] further reduces trainable parameters by using weight tying. Dora [<xref ref-type="bibr" rid="ref-38">38</xref>] enhances LoRA&#x2019;s learning ability and training stability by decomposing pre-trained weights into magnitude and direction. FedPara [<xref ref-type="bibr" rid="ref-39">39</xref>] improves fine-tuning efficiency by introducing Hadamard product reparameterization weights into the low-rank matrix, breaking the low-rank limitation.</p>
<p>Although Low-Rank Adaptation (LoRA) and its variants significantly reduce the computational cost of fine-tuning large pre-trained language models, their potential in multimodal fine-tuning remains largely unexplored. Our work investigates the application of LoRA in multimodal parameter-efficient transfer learning. By introducing trainable scaling vectors and cross-modal shared low-rank weight matrices, we achieve efficient knowledge transfer with minimal trainable parameters, while ensuring effective feature alignment between modalities.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>In this section, we first describe the framework of the multimodal foundational model we aim to fine-tune and the Low-Rank Adaptation (LoRA) [<xref ref-type="bibr" rid="ref-8">8</xref>]. We then introduce our parameter-efficient transfer learning method for multimodal models, UniTrans. This includes Vector-based Cross-modal Random Matrix Adaptation (VCRA), modality alignment design, as well as the Query-Augmentation Side Network (QASN) and Multimodal Consistency Alignment Module (MCAM).</p>
<sec id="s3_1">
<label>3.1</label>
<title>Preliminary</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Vision-Language Framework</title>
<p>We use BLIP as the backbone of our frozen pre-trained model. BLIP features a multimodal hybrid encoder-decoder structure (MED), unifying image-text matching and generation tasks. It employs a Vision Transformer (ViT) [<xref ref-type="bibr" rid="ref-40">40</xref>] as the image encoder and BERT [<xref ref-type="bibr" rid="ref-18">18</xref>] as the text encoder, with different components activated depending on the downstream multimodal task. For image-text matching tasks, cross-attention is added to the text encoder to fuse image and text features, and a special token <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">[</mml:mo><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is prepended to the input text. For generation tasks, a <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> token is inserted at the beginning of the input text, and the bidirectional self-attention layers are replaced with causal self-attention to generate captions for the given image.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Low-Rank Adaptation (LoRA)</title>
<p>LoRA utilizes low-rank matrices to approximate the weight changes during fine-tuning, effectively reducing the number of required parameters. Formally, for a pre-trained weight matrix <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, the weight update can be constrained as a low-rank matrix decomposition, as shown in <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. During fine-tuning, the original weights <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> remain frozen, and only the low-rank matrices are updated via gradient descent. Due to the low-rank nature, the dimension <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>r</mml:mi></mml:math></inline-formula> is typically small, making the size of the low-rank matrices significantly smaller than that of the original parameter matrix, where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>B</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>r</mml:mi><mml:mo>&#x226A;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. While LoRA offers an effective solution for efficient parameter transfer, it has not been further explored in the context of multimodal models.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Vector-Based Cross-Modal Random Matrix Adaptation (VCRA)</title>
<p>For pre-trained multimodal models, performing full fine-tuning on downstream tasks with small-scale datasets not only wastes computational resources but also risks knowledge forgetting and disrupting the feature alignment space. Therefore, we introduce additional trainable parameters to minimize or limit changes to the original parameters as much as possible. The model parameters updated through backpropagation using the fine-tuning data <italic>D</italic>:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>D</mml:mi><mml:mo>;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>W</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>LoRA adapts the weight space of the entire network by fine-tuning a matrix product of two low-rank matrices. However, directly applying LoRA to multimodal models yields unsatisfactory results due to the lack of interaction between modalities. To address this, we decompose the matrix into two low-rank matrices and two scaling vector as the projection, share knowledge across these matrices between modalities, and then apply projections to adapt the weight matrices of each layer for each modality.</p>
<p>Formally, compared to LoRA, VCRA not only decomposes the trainable weights into a pair of low-rank matrices <italic>A</italic> and <italic>B</italic>, but also introduces two trainable scaling vectors:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>W</mml:mi><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>B</mml:mi><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>W</mml:mi><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mi>B</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where trainable scaling vectors represented as diagonal matrices <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula>. These vectors effectively scale or deactivate specific rows and columns of the random matrices. The random matrices <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>A</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>B</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> in VCRA are shared across visual, textual and cross-modal modalities, where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>d</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>r</mml:mi></mml:math></inline-formula> represent the input and bottleneck dimensions, respectively. This sharing mechanism not only reduces the number of parameters significantly but also enhances cross-modal interaction:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>V</mml:mi><mml:mi>C</mml:mi><mml:mi>R</mml:mi><mml:mi>A</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>s</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msubsup><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mi>B</mml:mi><mml:mrow><mml:msubsup><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mi>A</mml:mi><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>s</mml:mi></mml:math></inline-formula> represents the learnable scaling factor, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>M</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, <italic>V</italic> denotes the visual modality, <italic>T</italic> denotes textual modality and <italic>C</italic> is cross-modal modality. Although we use shared random matrices for cross-modal information transfer, learning modality-specific knowledge is crucial for improving the transferability of multimodal models. Therefore, we apply modality-specific scaling vectors <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msubsup><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> to the visual, text encoders and feature fusion network with cross-attention layer, ensuring adaptability to each modality.</p>
<p>Compared to Lora, VeRA shares low rank matrices between different modalities and transformer layers and uses scaled vectors to adapt weight updates, greatly reducing the number of trainable parameters. Formally speaking, we use <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to denote the dimension of finetuned layers and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to represent the number of these layers. The number of trainable parameters for VCRA can be expressed as <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, contrasting with LoRA&#x2019;s <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mi>r</mml:mi></mml:math></inline-formula>, when we apply petl to the FFN of each layer. Specifically, for the lowest rank (i.e., <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>r</mml:mi></mml:math></inline-formula> &#x003D; 1), trainable parameters of VCRA is about half of LoRA. However, as <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>r</mml:mi></mml:math></inline-formula> insceases, the growth of LoRA trainable parameters is much faster than VCRA.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Modality Alignment Design</title>
<p>Besides VCRA, UniTrans also includes two additional modules, Query-Augmentation Side Network (QASN) and Multimodal Consistency Alignment Module (MCAM), to further enhance cross-modal alignment and multimodal fusion.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Query-Augmentation Side Network (QASN)</title>
<p>In the process of fine-tuning multimodal models, a factor that hinders modal alignment is that the multimodality fusion network is deep, which can lead to the loss of query information. To solve this problem, we designed a query information augmentation pipeline that runs parallel to the fusion network for adapted feature aggregation and information supplement.</p>
<p>Formally, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, given the multimodality fusion branch network consists of <italic>N</italic> transformer blocks, the forward process can be expressed as <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the <italic>N</italic>-th transformer block, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represent the image features and the text features, respectively, and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the fusion features. We apply VCRA to aggregate the fusion information from fusion branch network. Denote <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mi>B</mml:mi><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:mrow><mml:mi>A</mml:mi></mml:math></inline-formula> as the weight matrix accounting for the <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>i</mml:mi></mml:math></inline-formula>-th block, with <italic>A</italic> and <italic>B</italic> are the shared random matrices, and <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> are the scaling vectors. The query-augmentation side network gradually collects information from each block:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the output of the <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>h</mml:mi></mml:math></inline-formula>-th layer of QASN. At the same time, we will supplement the query information flowing in QSAN back into the fusion network:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Illustration of Vector-Based Cross-Modal Random Matrix Adaptation (shown above) and Query-Augmentation Side Network (shown below). VCRA achieves modal interaction through sharing low rank matrices and introduces trainable scaling vectors to adapt to weight updates of various modalities and layers. QASN enhances the fusion between modalities by adaptively supplementing query information in the fusion network</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-2.tif"/>
</fig>
<p>We did not take the query representation <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>h</mml:mi></mml:math></inline-formula> as the residual of the fusion representation <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, but instead train an information filter to explicitly model the contribution of supplementary query information. Our filter design is simple and effective, and can be represented as <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi mathvariant="bold-italic">&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">&#x03B2;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Here, <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> represent learnable parameter matrices, which are initialized to <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mn>0</mml:mn></mml:math></inline-formula>, respectively. Finally, the filtered information is processed through the activation function <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi mathvariant="bold-italic">&#x03C3;</mml:mi></mml:math></inline-formula>. As a result, multimodality fusion branch network with supplementary query information is more conducive to transfer knowledge for downstream tasks.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Multimodal Consistency Alignment Module (MCAM)</title>
<p>In addition to using unsupervised contrastive learning during pre-training, BLIP also employs it during downstream task fine-tuning. As we know, with sufficient training samples, the contrastive loss function effectively brings positive pairs closer and pushes negative pairs farther apart. However, the scale of downstream task fine-tuning datasets is typically limited, which allows individual noisy samples to interfere with the feature representation process, potentially disrupting the previously aligned feature space. To address this issue, we introduce MCAM (<xref ref-type="fig" rid="fig-3">Fig. 3</xref>), which captures fine-grained relationships between samples through rank consistency, mitigating the impact of noisy samples.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Illustration of Multimodal Consistency Alignment Module, it enhances fine-grained alignment between modalities by matching the similarity ranking consistency between image-text pairs in the contrastive learning process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-3.tif"/>
</fig>
<p>Using contrastive learning loss to align image and text features lacks modeling the degree of repulsion for negative samples. For example, pushing an picture of a dog &#x2019;equally&#x2019; away from the descriptions &#x201D;cat&#x201D;, &#x201D;tiger&#x201D;, and &#x201D;house&#x201D; can lead to the loss of fine-grained similarity information between the image and text and affecting feature alignment. We introduce the Multimodal Consistency Alignment Module to solve this problem by matching the consistency of similarity rankings between image-text pairs.</p>
<p>Formally, for a mini-batch of image-text pairs denoted as <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the image and <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the corresponding text, image and text are processed through the visual encoder <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and text encoder <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively, to obtain modality-specific representations. BLIP uses <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as the contrastive loss:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>ITC</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mspace width="2em" /><mml:mspace width="1em" /><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> is a temperature hyperparameter. While <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is effective at distinguishing between positive and negative image-text pairs, it does not consider differences between highly relevant and moderately relevant pairs. Our proposed multimodal consistency alignment module introduces ranking information to capture fine-grained image-text relationships, enhancing modality representation and strengthening cross-modal alignment.</p>
<p>Specifically, for a given image-text pair within a mini-batch <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, we can obtain a list <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="bold-italic">&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> that represents the cosine similarity between the image <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and all texts in the batch, as well as a list <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi mathvariant="bold-italic">&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> that represents the cosine similarity between the text <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and all images in the batch. We aim for the corresponding elements in <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to have the same ranking positions. This allows us to capture the fine-grained ranking information between the image and various negative text samples, as well as between the text and different negative image samples. We achieve ranking consistency by minimizing the the Jensen-Shannon (JS) divergence of the two top one probability distributions:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>ITR</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mtext>JS</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext>S</mml:mtext></mml:mrow><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mspace width="1em" /><mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo><mml:mo>+</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mrow><mml:mover><mml:mi>S</mml:mi><mml:mo>&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represent the top-one probability distributions of <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively, and <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> is a hyperparameter. The final loss function is:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>all</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>ITC</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>ITR</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Implementation Details</title>
<p>We apply BLIP-base as our vision-language backbone for Image Caption, Image-Text Retrieval and VQA downstream tasks. During the fine-tuning process, the parameters of the backbone model are kept frozen. The experiments were implemented using PyTorch on 8 <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> NVIDIA 3090 GPU. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, we present the training details for UniTrans. We applied VCRA to the projection of query, key and value in Attention layer, and the scaling vectors of key projection share parameters with value projection. At the start of fine-tuning, we initialize the random matrices <italic>A</italic> and <italic>B</italic> with random values drawn from a normal distribution, while vector <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>a</mml:mi></mml:msub></mml:math></inline-formula> is initialized to <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mn>1</mml:mn></mml:math></inline-formula> and vector <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi mathvariant="normal">&#x039B;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math></inline-formula> is initialized to <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mn>0</mml:mn></mml:math></inline-formula>. For the query information filters in QASN (components <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula>), are initialized to <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mn>1</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mn>0</mml:mn></mml:math></inline-formula>, respectively. For the textual data in video datasets, we performed simple preprocessing steps, such as truncating words that exceed the maximum sentence length.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Setting hyperparameters for fine-tuning training of multimodal downstream tasks</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Config</th>
<th align="center">Image captioning</th>
<th align="center" colspan="3">Image-text retrieval</th>
<th align="center" colspan="2">Visual question answering</th>
</tr>
<tr align="center">
<th>COCO (caption)</th>
<th>MSCOCO</th>
<th>Flickr</th>
<th>DiDemo</th>
<th>VQAv2</th>
<th>MSRVTT-QA</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td>learning rate</td>
<td>1e&#x2212;5</td>
<td>1e&#x2212;5</td>
<td>1e&#x2212;4</td>
<td>1e&#x2212;5</td>
<td>2e&#x2212;5</td>
<td>2e&#x2212;5</td>
</tr>
<tr align="center">
<td>batch size</td>
<td>128</td>
<td>128</td>
<td>128</td>
<td>32</td>
<td>128</td>
<td>64</td>
</tr>
<tr align="center">
<td>epochs</td>
<td>6</td>
<td>5</td>
<td>6</td>
<td>10</td>
<td>10</td>
<td>10</td>
</tr>
<tr align="center">
<td>training input</td>
<td>384</td>
<td>384</td>
<td>384</td>
<td>8 <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224</td>
<td>384</td>
<td>8 <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224</td>
</tr>
<tr align="center">
<td>inference input</td>
<td>384</td>
<td>384</td>
<td>384</td>
<td>16 <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224</td>
<td>384</td>
<td>16 <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Baselines &#x00026; Datasets &#x00026; Evaluation Metrics</title>
<p>We evaluate UniTrans across six benchmarks, covering three cross-modal tasks: Image Captioning, Image-Text Retrieval, and VQA. For the image captioning task, we use the COCO-Caption [<xref ref-type="bibr" rid="ref-41">41</xref>] dataset with the COCO Caption Karpathy split as the test set, employing BLEU@4 and CIDEr as evaluation metrics. BLEU measures the n-gram precision between the generated and reference captions, while CIDEr evaluates the similarity between candidate and reference captions by calculating the cosine similarity of their TF-IDF vectors. For the image-text retrieval task, we use the MSCOCO [<xref ref-type="bibr" rid="ref-41">41</xref>], Flickr30K [<xref ref-type="bibr" rid="ref-42">42</xref>], and Didemo datasets [<xref ref-type="bibr" rid="ref-43">43</xref>], with Recall at K (R@K) as evaluation metrics. R@K aims to calculate the ratio of queries that successfully retrieve the ground truth as one of the first K results. For the VQA task, we use the VQAv2 [<xref ref-type="bibr" rid="ref-44">44</xref>] and MSRVTT-QA [<xref ref-type="bibr" rid="ref-45">45</xref>] datasets, with the evaluation results obtained from the official validation platforms provided by the datasets.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Results</title>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Performance Comparisons on Cross-modal Tasks</title>
<p><xref ref-type="table" rid="table-2">Tables 2</xref> and <xref ref-type="table" rid="table-3">3</xref> show the performances of UniTrans for image-text retrieval task on Flickr30K and MSCOCO. As shown, UniTrans achieves performance comparable to UniAdapter with only 0.2M parameters on both Flickr30K and MSCOCO, even outperforming it on certain metrics. Additionally, our method&#x2019;s performance is very close to that of fully fine-tuning the BLIP backbone while exceeding previous fully fine-tuned methods.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Results on flickr</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="3">Text retrieval</th>
<th colspan="3">Image retrieval</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td colspan="8"><bold>Full fine-tuning</bold></td>
</tr>
<tr align="center">
<td>UNITER [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>330M</td>
<td>87.3</td>
<td>98.0</td>
<td>99.2</td>
<td>75.6</td>
<td>94.1</td>
<td>96.8</td>
</tr>
<tr align="center">
<td>UNIMO [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>330M</td>
<td>89.4</td>
<td>98.9</td>
<td>99.8</td>
<td>78.0</td>
<td>94.2</td>
<td>97.1</td>
</tr>
<tr align="center">
<td>ALIGN [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>820M</td>
<td>95.3</td>
<td>99.8</td>
<td><bold>100.0</bold></td>
<td>84.9</td>
<td>97.4</td>
<td>98.6</td>
</tr>
<tr align="center">
<td>ALBEF [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>210M</td>
<td>95.9</td>
<td>99.8</td>
<td><bold>100.0</bold></td>
<td>85.6</td>
<td>97.5</td>
<td><bold>98.9</bold></td>
</tr>
<tr align="center">
<td>BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>223M</td>
<td><bold>97.3</bold></td>
<td>99.9</td>
<td><bold>100.0</bold></td>
<td><bold>87.3</bold></td>
<td><bold>97.6</bold></td>
<td><bold>98.9</bold></td>
</tr>
<tr align="center">
<td colspan="8"><bold>Frozen backbone</bold></td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>10.6M</td>
<td>96.2</td>
<td>99.7</td>
<td>99.8</td>
<td>85.8</td>
<td>97.1</td>
<td>98.4</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>97.1</td>
<td><bold>100.0</bold></td>
<td><bold>100.0</bold></td>
<td>86.5</td>
<td>97.4</td>
<td>98.8</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>97.1</td>
<td>99.9</td>
<td><bold>100.0</bold></td>
<td>86.4</td>
<td>97.4</td>
<td><bold>98.9</bold></td>
</tr>
<tr align="center">
<td>UniTrans (ours, r &#x003D; 64)</td>
<td><bold>0.2M</bold></td>
<td>97.2</td>
<td><bold>100.0</bold></td>
<td><bold>100.0</bold></td>
<td>86.4</td>
<td>97.4</td>
<td>98.8</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Results on MSCOCO</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="3">Text retrieval</th>
<th colspan="3">Image retrieval</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td align="center" colspan="8"><bold>Full fine-tuning</bold></td>
</tr>
<tr align="center">
<td>Unicoder-VL [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>&#x2013;</td>
<td>62.3</td>
<td>87.1</td>
<td>92.8</td>
<td>46.7</td>
<td>76.0</td>
<td>85.3</td>
</tr>
<tr align="center">
<td>OSCAR [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>330M</td>
<td>70.0</td>
<td>91.1</td>
<td>95.5</td>
<td>54.0</td>
<td>80.8</td>
<td>88.5</td>
</tr>
<tr align="center">
<td>ALIGN [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>820M</td>
<td>77.0</td>
<td>93.5</td>
<td>96.9</td>
<td>59.9</td>
<td>83.3</td>
<td>89.8</td>
</tr>
<tr align="center">
<td>ALBEF [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>210M</td>
<td>77.6</td>
<td>94.3</td>
<td>97.2</td>
<td>60.7</td>
<td>84.3</td>
<td>90.5</td>
</tr>
<tr align="center">
<td>BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>223M</td>
<td><bold>81.9</bold></td>
<td><bold>95.4</bold></td>
<td><bold>97.8</bold></td>
<td><bold>64.3</bold></td>
<td><bold>85.7</bold></td>
<td><bold>91.5</bold></td>
</tr>
<tr align="center">
<td align="center" colspan="8"><bold>Frozen backbone</bold></td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>10.6M</td>
<td>80.0</td>
<td>94.1</td>
<td>97.2</td>
<td>62.1</td>
<td>84.4</td>
<td>90.6</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>79.8</td>
<td>94.2</td>
<td>97.5</td>
<td>62.3</td>
<td>84.5</td>
<td>90.8</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>80.1</td>
<td>94.6</td>
<td>97.4</td>
<td>62.6</td>
<td>84.6</td>
<td>90.9</td>
</tr>
<tr align="center">
<td>UniTrans (ours, r &#x003D; 64)</td>
<td><bold>0.2M</bold></td>
<td>79.7</td>
<td>94.2</td>
<td>97.4</td>
<td>62.2</td>
<td>84.4</td>
<td>90.5</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-3fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Unlike image-text retrieval, the VQA task requires multimodal generation capabilities. Therefore, in addition to the encoder, a decoder is necessary for text generation. Due to the structural differences between the encoder and decoder, we did not share the random matrices parameters, resulting in a slight increase in trainable parameters. However, the total parameter count remains significantly lower than the baseline. According to <xref ref-type="table" rid="table-4">Table 4</xref>, our method outperforms all fine-tuning methods, demonstrating that our approach is well-suited for multimodal generation tasks.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Results on VQAv2</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="2">VQAv2</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>Test-dev</th>
<th>Test-std</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td align="center" colspan="4"><bold>Full fine-tuning</bold></td>
</tr>
<tr align="center">
<td>VL-T5/BART [<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
<td>165M</td>
<td>&#x2013;</td>
<td>71.30</td>
</tr>
<tr align="center">
<td>SOHO [<xref ref-type="bibr" rid="ref-52">52</xref>]</td>
<td>155M</td>
<td>73.25</td>
<td>73.47</td>
</tr>
<tr align="center">
<td>OSCAR [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>330M</td>
<td>73.61</td>
<td>73.82</td>
</tr>
<tr align="center">
<td>UNITER [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>330M</td>
<td>73.82</td>
<td>74.03</td>
</tr>
<tr align="center">
<td>ALBEF [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>266M</td>
<td>75.84</td>
<td>76.04</td>
</tr>
<tr align="center">
<td>BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>337M</td>
<td>77.44</td>
<td>77.48</td>
</tr>
<tr align="center">
<td align="center" colspan="4"><bold>Frozen backbone</bold></td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>73.72</td>
<td>73.71</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>75.44</td>
<td>75.56</td>
</tr>
<tr align="center">
<td>UniTrans (ours, r &#x003D; 64)</td>
<td><bold>0.3M</bold></td>
<td><bold>77.68</bold></td>
<td><bold>77.77</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-4fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows the performance of our method on the Image Captioning task using the COCO caption dataset. The results indicate that UniTrans outperforms the baseline and is only slightly below the fully fine-tuned BLIP method. As shown in <xref ref-type="table" rid="table-6">Table 6</xref>, our UniTrans also outperforms the baseline on video datasets.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Results on the Image Captioning dataset COCO Caption, where B@4 represents BLEU@4 and C denotes CIDEr</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Pre-train</th>
<th colspan="2">COCO caption karpathy test</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>B@4</th>
<th>C</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td colspan="4"><bold>Full fine-tuning</bold></td>
</tr>
<tr align="center">
<td>Enc-Dec [<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>15M</td>
<td>&#x2013;</td>
<td>110.9</td>
</tr>
<tr align="center">
<td>VinVL [<xref ref-type="bibr" rid="ref-54">54</xref>]</td>
<td>5.7M</td>
<td>38.2</td>
<td>129.3</td>
</tr>
<tr align="center">
<td>LEMON [<xref ref-type="bibr" rid="ref-55">55</xref>]</td>
<td>200M</td>
<td><bold>40.3</bold></td>
<td><bold>133.3</bold></td>
</tr>
<tr align="center">
<td>BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>14M</td>
<td>38.6</td>
<td>129.7</td>
</tr>
<tr align="center">
<td>BLIP [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>129M</td>
<td>39.7</td>
<td><bold>133.3</bold></td>
</tr>
<tr align="center">
<td colspan="4"><bold>Frozen backbone</bold></td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 64)</td>
<td>129M</td>
<td>38.88</td>
<td>131.5</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>129M</td>
<td>39.0</td>
<td>132.1</td>
</tr>
<tr align="center">
<td>UnTrans (ours, r &#x003D; 64)</td>
<td>129M</td>
<td>39.2</td>
<td>132.3</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Results on the video-text retrieval dataset DiDemo and the video visual question answering dataset MSRVTT-QA</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col align="center" width="35mm"/>
<col align="center" width="15mm"/>
<col align="center" width="20mm"/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="3">DiDemo</th>
<th>Method</th>
<th># Tunable</th>
<th>MSRVTT-QA</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
<th></th>
<th></th>
<th>Test acc</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td colspan="8"><bold>Full fine-tuning</bold></td>
</tr>
<tr align="center">
<td>CLIPBERT [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>135M</td>
<td>20.4</td>
<td>48.0</td>
<td>60.8</td>
<td>CLIPBERT [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>135M</td>
<td>37.4</td>
</tr>
<tr align="center">
<td>Frozen in Time [<xref ref-type="bibr" rid="ref-57">57</xref>]</td>
<td>180M</td>
<td>34.6</td>
<td>65.0</td>
<td>74.7</td>
<td>CoMVT [<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>&#x2013;</td>
<td>39.5</td>
</tr>
<tr align="center">
<td>ALPRO [<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>245M</td>
<td>35.9</td>
<td>67.5</td>
<td>78.8</td>
<td>ALPRO [<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>245M</td>
<td>42.1</td>
</tr>
<tr align="center">
<td>VIOLET [<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>306M</td>
<td>32.6</td>
<td>62.8</td>
<td>74.7</td>
<td>Just-Ask [<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>200M</td>
<td>41.5</td>
</tr>
<tr align="center">
<td>All-in-one [<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
<td>110M</td>
<td>32.7</td>
<td>61.4</td>
<td>73.5</td>
<td>VIOLET [<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>306M</td>
<td>43.9</td>
</tr>
<tr align="center">
<td>CLIP4Clip [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>124M</td>
<td>42.8</td>
<td>68.5</td>
<td>79.2</td>
<td>MERLOT [<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
<td>233M</td>
<td>43.1</td>
</tr>
<tr align="center">
<td colspan="8"><bold>Frozen backbone</bold></td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>10.6M</td>
<td>50.9</td>
<td>75.3</td>
<td>82.4</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>49.0</td>
<td>75.5</td>
<td>83.3</td>
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>44.2</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>52.1</td>
<td><bold>77.3</bold></td>
<td>85.2</td>
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>44.7</td>
</tr>
<tr align="center">
<td>UniTrans (ours, r &#x003D; 64)</td>
<td><bold>0.2M</bold></td>
<td><bold>52.4</bold></td>
<td><bold>77.3</bold></td>
<td><bold>85.3</bold></td>
<td>UniTrans (ours, r &#x003D; 64)</td>
<td><bold>0.3M</bold></td>
<td><bold>44.8</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-6fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>From the comparative experimental results on multimodal benchmarks, it can be observed that Unitrans achieves competitive performance while reducing the trainable parameters of LoRA to just 5%. This is attributable to two main factors. First, we leverage VeRA to share a low-rank random matrix across modalities and employ scaling vectors to adapt weight updates. This shared mechanism reduces the number of large random matrices involved in training, effectively decreasing parameter size while enhancing inter-modal information exchange. Second, QASN adaptively supplements query information within the fusion network, and MCAM regularizes fine-grained image-text alignment during fine-tuning through contrastive learning loss, strengthening multimodal alignment. In comparison, UniAdapter enables modality interaction during fine-tuning by employing a shared MLP in the adapter. However, it lacks information sharing across layers, and the MLP itself involves a considerable number of trainable parameters, particularly as the rank <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>r</mml:mi></mml:math></inline-formula> increases. Based on this, we extend the concept of UniAdapter to Low-Rank Adaptation and further optimize it by introducing QASN and MCAM to facilitate multimodal alignment during fine-tuning. As a result, we achieve performance comparable to UniAdapter while significantly reducing the number of trainable parameters.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Training Efficiency and Storage Cost</title>
<p>As shown in <xref ref-type="table" rid="table-7">Table 7</xref>, we present the training time and GPU memory costs for the three tasks: Image Captioning, Image-Text Retrieval, and Visual Question Answering. We consider the training time and storage cost of fully fine-tuning BLIP as one unit. From the table, it can be observed that our UniTrans outperforms both fully fine-tuned BLIP and UniAdapter in terms of training time and GPU memory cost.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Comparison of training time and GPU memory usage</title>
</caption>
<table>
<colgroup>
<col width="30mm"/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th>#Tunable</th>
<th colspan="2">Image captioning</th>
<th>#Tunable</th>
<th colspan="2">Image-text retrieval</th>
<th>#Tunable</th>
<th colspan="2">VQA</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>Time</th>
<th>Memory</th>
<th></th>
<th>Time</th>
<th>Memory</th>
<th></th>
<th>Time</th>
<th>Memory</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td>Full fine-tuning</td>
<td>213M</td>
<td>1.00</td>
<td>1.00</td>
<td>223M</td>
<td>1.00</td>
<td>1.00</td>
<td>337M</td>
<td>1.00</td>
<td>1.00</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 512)</td>
<td>19.0M</td>
<td>0.83</td>
<td>0.81</td>
<td>19.0M</td>
<td>0.88</td>
<td>0.86</td>
<td>19.0M</td>
<td>0.93</td>
<td>0.80</td>
</tr>
<tr align="center">
<td>UniAdapter (r &#x003D; 128)</td>
<td>4.8M</td>
<td>0.81</td>
<td>0.77</td>
<td>4.8M</td>
<td>0.86</td>
<td>0.82</td>
<td>4.8M</td>
<td>0.92</td>
<td><bold>0.72</bold></td>
</tr>
<tr align="center">
<td>UniTrans (r &#x003D; 128)</td>
<td>0.36M</td>
<td>0.81</td>
<td>0.77</td>
<td>0.36M</td>
<td>0.86</td>
<td>0.81</td>
<td>0.46M</td>
<td>0.89</td>
<td><bold>0.71</bold></td>
</tr>
<tr align="center">
<td>UniTrans (r &#x003D; 64)</td>
<td><bold>0.26M</bold></td>
<td><bold>0.80</bold></td>
<td><bold>0.75</bold></td>
<td><bold>0.26M</bold></td>
<td><bold>0.83</bold></td>
<td><bold>0.79</bold></td>
<td><bold>0.35M</bold></td>
<td><bold>0.89</bold></td>
<td><bold>0.68</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-8fn1" fn-type="other">
<p>Note: Bold represents optimal performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>The Impact of Rank on UniTrans</title>
<p>We also explored the impact of rank size on UniTrans. As shown in<xref ref-type="table" rid="table-4"> Fig. 4a</xref>,<xref ref-type="table" rid="table-4">b</xref>, experiments conducted on Flickr30K indicate that as the rank increases, performance of cross-modal downstream fine-tuning improves. However, when the rank reaches a certain threshold, the performance gains slow down or even slightly decrease. This may be due to the rank increase making the random matrix too large, introducing redundant parameters and leading to overfitting. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4c</xref>, as the rank R increases, the growth in parameters for our UniTrans is significantly slower compared to LoRA and UniAdapter, demonstrating the superiority of our method.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>how rank affects UniTrans. (a) and (b) show the impact of different ranks on UniTrans performance on the Flickr30K dataset. (c) demonstrates the scalability of parameters compared to other PETL methods</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-4.tif"/>
</fig>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Visualization Results</title>
<p>To intuitively understand the performance of UniTrans in multimodal downstream tasks, we present detailed visualization results in <xref ref-type="fig" rid="fig-5">Figs. 5</xref>&#x2013;<xref ref-type="fig" rid="fig-7">7</xref>. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows a comparison between the captions generated by UniTrans on the COCO Caption dataset and the ground-truth. It can be seen that UniTrans is capable of accurately generating image captions, with key words closely matching the ground-truth. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> demonstrates the results of the Image-Text Retrieval task on the Flickr30K dataset, where UniTrans accurately retrieves the corresponding image given a textual caption. We also present the results of the VQA task on the VQAv2 dataset, where UniTrans provides accurate answers. However, the responses are somewhat brief, which may be attributed to the relatively smaller parameter size of the multimodal base model. In the future, we plan to conduct experiments with larger-scale models.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Visualization results on COCO Caption</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Visualization results on Flickr30K</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-6.tif"/>
</fig><fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Visualization results on VQAv2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_59745-fig-7.tif"/>
</fig>
</sec>
<sec id="s4_3_5">
<label>4.3.5</label>
<title>Ablation Study</title>
<p>To assess the effectiveness of each module in UniTrans, we conducted ablation experiments on image-text retrieval using the Flickr30K benchmark and the DiDemo benchmark. We compared UniTrans with three baselines: a frozen BLIP, fully fine-tuned BLIP, and LoRA. As shown in <xref ref-type="table" rid="table-8">Tables 8</xref> and <xref ref-type="table" rid="table-9">9</xref>, when the random matrices are not shared (&#x2018;&#x002B;scaling vectors&#x2019; means only scaling vector is used, and random matrices are shared in each transformer block of single modality), the number of trainable parameters is significantly reduced, but the performance deteriorates. However, our proposed VCRA demonstrates clear improvements, highlighting its contribution to modality alignment. The results also indicate that the QASN and the MCAM, both designed to enhance modality alignment, further boost performance. Notably, the latter achieves performance gains without introducing additional parameters.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Results of ablation on the Flickr</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="3">Text retrieval</th>
<th colspan="3">Image retrieval</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td>Frozen</td>
<td>0M</td>
<td>86.9</td>
<td>98.0</td>
<td>99.1</td>
<td>78.1</td>
<td>94.0</td>
<td>97.7</td>
</tr>
<tr align="center">
<td>Full fine-tuning</td>
<td>223M</td>
<td>97.3</td>
<td>99.9</td>
<td>100.0</td>
<td>87.3</td>
<td>97.6</td>
<td>98.9</td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>10.6M</td>
<td>96.2</td>
<td>99.7</td>
<td>99.8</td>
<td>85.8</td>
<td>97.1</td>
<td>98.4</td>
</tr>
<tr align="center">
<td>&#x002B;scaling vectors</td>
<td>0.28M</td>
<td>92.0</td>
<td>98.6</td>
<td>99.5</td>
<td>82.1</td>
<td>95.9</td>
<td>98.0</td>
</tr>
<tr align="center">
<td>&#x002B;shared random matrices</td>
<td>0.18M</td>
<td>95.3</td>
<td>99.6</td>
<td>99.9</td>
<td>84.4</td>
<td>96.7</td>
<td>98.4</td>
</tr>
<tr align="center">
<td>&#x002B;QASN</td>
<td>0.21M</td>
<td>96.1</td>
<td>99.7</td>
<td>100.0</td>
<td>85.2</td>
<td>96.8</td>
<td>98.6</td>
</tr>
<tr align="center">
<td>&#x002B;MCAM (UniTrans)</td>
<td>0.21M</td>
<td>96.5</td>
<td>99.8</td>
<td>100.0</td>
<td>86.0</td>
<td>97.2</td>
<td>98.8</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Results of ablation on the DiDemo</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr align="center">
<th>Method</th>
<th># Tunable</th>
<th colspan="3">DiDemo</th>
</tr>
<tr align="center">
<th></th>
<th></th>
<th>R@1</th>
<th>R@5</th>
<th>R@10</th>
</tr>
</thead>
<tbody>
<tr align="center">
<td>Linear probe</td>
<td>0.4M</td>
<td>39.7</td>
<td>64.6</td>
<td>74.9</td>
</tr>
<tr align="center">
<td>Full fine-tuning</td>
<td>223M</td>
<td>51.3</td>
<td>79.1</td>
<td>85.7</td>
</tr>
<tr align="center">
<td>LoRA (r &#x003D; 32)</td>
<td>10.6M</td>
<td>50.9</td>
<td>75.3</td>
<td>82.4</td>
</tr>
<tr align="center">
<td>&#x002B;scaling vectors</td>
<td>0.28M</td>
<td>48.4</td>
<td>74.9</td>
<td>81.6</td>
</tr>
<tr align="center">
<td>&#x002B;shared random matrices</td>
<td>0.18M</td>
<td>49.1</td>
<td>76.2</td>
<td>83.3</td>
</tr>
<tr align="center">
<td>&#x002B;QASN</td>
<td>0.21M</td>
<td>50.6</td>
<td>76.5</td>
<td>84.0</td>
</tr>
<tr align="center">
<td>&#x002B;MCAM (UniTrans)</td>
<td>0.21M</td>
<td>51.8</td>
<td>77.1</td>
<td>85.2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion &#x0026; Future Work</title>
<p>In this paper, we propose a novel paradigm for efficient cross-modal knowledge transfer, UniTrans, to enhance modality interaction and alignment during the fine-tuning process of multimodal models. The concepts of VCRA, along with the modality alignment modules QASN and MCAM, are simple and lightweight, allowing them to be extended to different multimodal base models without altering their inherent structure, thereby effectively adapting to various fine-grained visual-language tasks. Extensive evaluations on multiple downstream benchmarks demonstrate that our method achieves superior performance with fewer than 1M parameters.</p>
<p>Moreover, UniTrans has its limitations, which provide directions for our future work. (1) UniTrans has only been validated on the multimodal model architecture with fusion networks represented by BLIP, and has not been tested on architectures such as the Q-Former (represented by BLIP2) or the MLP-based architecture (represented by LLava). In the future, we will explore a wider range of multimodal model architectures. (2) Our method has been compared to classic cross-modal tasks such as Image Captioning, Image-Text Retrieval, and VQA, but has not yet been explored for other cross-modal tasks, such as text-to-image generation. Moving forward, we will extend our approach to a broader set of tasks. (3) The fine-tuning datasets we used are relatively small in scale. In the future, we plan to conduct experiments on larger-scale datasets.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Jiakang Sun, Ke Chen; data collection: Xinyang He; analysis and interpretation of results: Jiakang Sun, Xu Liu, Ke Li; draft manuscript preparation: Jiakang Sun, Cheng Peng. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available in the The website link below: MSCOCO: <ext-link ext-link-type="uri" xlink:href="https://cocodataset.org/#download">https://cocodataset.org/#download</ext-link>, accessed on 10 July 2024. Flickr30K: <ext-link ext-link-type="uri" xlink:href="https://shannon.cs.illinois.edu/DenotationGraph/data/index.html">https://shannon.cs.illinois.edu/DenotationGraph/data/index.html</ext-link>, accessed on 10 July 2024. MSRVTT: <ext-link ext-link-type="uri" xlink:href="https://www.mediafire.com/folder/h14iarbs62e7p/shared">https://www.mediafire.com/folder/h14iarbs62e7p/shared</ext-link>, accessed on 10 July 2024. DiDemo: <ext-link ext-link-type="uri" xlink:href="https://github.com/jpthu17/EMCL">https://github.com/jpthu17/EMCL</ext-link>, accessed on 10 July 2024. VQAv2: <ext-link ext-link-type="uri" xlink:href="https://visualqa.org/download.html">https://visualqa.org/download.html</ext-link>, accessed on 10 July 2024. COCO Caption: https://cocodataset.org/#download, accessed on 10 July 2024.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Raji&#x010D;</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tai</surname> <given-names>YW</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>CK</given-names></string-name>, <string-name><surname>Danelljan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Segment anything meets point tracking</article-title>. <comment>arXiv:230701197. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2307.01197</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Touvron</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lavril</surname> <given-names>T</given-names></string-name>, <string-name><surname>Izacard</surname> <given-names>G</given-names></string-name>, <string-name><surname>Martinet</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lachaux</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Lacroix</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LLaMA: open and efficient foundation language models</article-title>. <comment>arXiv:230213971. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2302.13971</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kirillov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mintun</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ravi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Rolland</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gustafson</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Segment anything</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <year>2023</year>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>4015</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV51070.2023.00371</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2022</year>; <conf-loc>Baltimore, MD, USA: PMLR.</conf-loc> p. <fpage>12888</fpage>&#x2013;<lpage>900</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Brown</surname> <given-names>TB</given-names></string-name></person-group>. <article-title>Language models are few-shot learners</article-title>. <comment>arXiv:2005.14165</comment>. <year>2020</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2005.14165</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kheddar</surname> <given-names>H</given-names></string-name>, <string-name><surname>Himeur</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Al-Maadeed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Amira</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bensaali</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Deep transfer learning for automatic speech recognition: towards better generalization</article-title>. <source>Knowl Based Syst</source>. <year>2023</year>;<volume>277</volume>:<fpage>110851</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.knosys.2023.110851</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Arabnia</surname> <given-names>HR</given-names></string-name>, <string-name><surname>Rasheed</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A review of deep transfer learning and recent advancements</article-title>. <source>Technologies</source>. <year>2023</year>;<volume>11</volume>(<issue>2</issue>):<fpage>40</fpage>. doi:<pub-id pub-id-type="doi">10.3390/technologies11020040</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>EJ</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wallis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Allen-Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LoRA: low-rank adaptation of large language models</article-title>. <comment>arXiv:2106.09685. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2106.09685</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Houlsby</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giurgiu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jastrzebski</surname> <given-names>S</given-names></string-name>, <string-name><surname>Morrone</surname> <given-names>B</given-names></string-name>, <string-name><surname>De Laroussilhe</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Gesmundo</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Parameter-efficient transfer learning for NLP</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2019</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>. p. <fpage>2790</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>XL</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Prefix-tuning: optimizing continuous prompts for generation</article-title>. <comment>arXiv:2101.00190</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2101.00190</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cho</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Arnab</surname> <given-names>A</given-names></string-name>, <string-name><surname>Seo</surname> <given-names>PH</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Cost aggregation for open-vocabulary semantic segmentation</article-title>. In: <conf-name>CAT-seg: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2024</year>; <publisher-loc>Seattle WA, USA</publisher-loc>. p. <fpage>4113</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52733.2024.00394</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>G</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DenseCLIP: language-guided dense prediction with context-aware prompting</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2022</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>18082</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01755</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sung</surname> <given-names>YL</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Vl-adapter: parameter-efficient transfer learning for vision-and-language tasks</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2022</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>5227</fpage>&#x2013;<lpage>37</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tomizuka</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>UniAdapter: unified parameter-efficient transfer learning for cross-modal modeling</article-title>. <comment>arXiv:2302.06605. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2302.06605</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning transferable visual models from natural language supervision</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2021</year>; <publisher-loc>PMLR</publisher-loc>. p. <fpage>8748</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Segment anything in 3D with NeRFs</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2023</year>;<volume>36</volume>:<fpage>25971</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kheddar</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Transformers and large language models for efficient intrusion detection systems: a comprehensive survey</article-title>. <comment>arXiv:2408.07583. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2408.07583</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name></person-group>. <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. <comment>arXiv:1810.04805. 2018</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1810.04805</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Savarese</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2023</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>: <publisher-loc>PMLR</publisher-loc>. p. <fpage>19730</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>YJ</given-names></string-name></person-group>. <article-title>Visual instruction tuning</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2024</year>;<volume>36</volume>:<fpage>34892</fpage>&#x2013;<lpage>916</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Han</surname> <given-names>J</given-names></string-name></person-group>. <article-title>IMRAM: iterative matching with recurrent attention memory for cross-modal image-text retrieval</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2020</year>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>12655</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01267</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Qu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Dynamic modality interaction modeling for image-text retrieval</article-title>. In: <conf-name>Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>; <year>2021</year>; <publisher-loc>Canada</publisher-loc>. p. <fpage>1104</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3404835.3462829</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Smola</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Stacked attention networks for image question answering</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2016</year>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>21</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2016.10</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Unicoder-VL: a universal encoder for vision and language by cross-modal pre-training</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>:<fpage>11336</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i07.6795</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gai</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhuang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>G</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Context-I2W: mapping images to context-dependent words for accurate zero-shot composed image retrieval</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2024</year>;<volume>38</volume>:<fpage>5180</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i6.28324</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>LocVTP: video-text pre-training for temporal localization</article-title>. In: <conf-name>European Conference on Computer Vision</conf-name>; <year>2022</year>; <publisher-loc>Tel Aviv, Israel</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>38</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-19809-0_3</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhuang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Towards fast and accurate image-text retrieval with self-supervised fine-grained alignment</article-title>. <source>IEEE Trans Multimed</source>. <year>2023</year>;<volume>26</volume>:<fpage>1361</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TMM.2023.3280734</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Dense contrastive learning for self-supervised visual pre-training</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2021</year>. p. <fpage>3024</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00304</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gai</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Align before Search: aligning ads image to text for accurate cross-modal sponsored search</article-title>. <comment>arXiv:2309.16141. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2309.16141</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lester</surname> <given-names>B</given-names></string-name>, <string-name><surname>Al-Rfou</surname> <given-names>R</given-names></string-name>, <string-name><surname>Constant</surname> <given-names>N</given-names></string-name></person-group>. <article-title>The power of scale for parameter-efficient prompt tuning</article-title>. <comment>arXiv:2104.08691. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2104.08691</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>GPT understands, too</article-title>. <source>AI Open</source>. <year>2024</year>;<volume>5</volume>:<fpage>208</fpage>&#x2013;<lpage>215</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aiopen.2023.08.012</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pfeiffer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kamath</surname> <given-names>A</given-names></string-name>, <string-name><surname>R&#x00FC;ckl&#x00E9;</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gurevych</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Adapterfusion: non-destructive task composition for transfer learning</article-title>. <comment>arXiv:2005.00247. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2005.00247</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>R&#x00FC;ckl&#x00E9;</surname> <given-names>A</given-names></string-name>, <string-name><surname>Geigle</surname> <given-names>G</given-names></string-name>, <string-name><surname>Glockner</surname> <given-names>M</given-names></string-name>, <string-name><surname>Beck</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pfeiffer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Reimers</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AdapterDrop: on the efficiency of adapters in transformers</article-title>. <comment>arXiv:2010.11918. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2010.11918</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zaken</surname> <given-names>EB</given-names></string-name>, <string-name><surname>Ravfogel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Goldberg</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>BitFit: simple parameter-efficient fine-tuning for transformer-based masked language-models</article-title>. <comment>arXiv:2106.10199. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2106.10199</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Rush</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Parameter-efficient transfer learning with diff pruning</article-title>. <comment>arXiv:2012.07463. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2012.07463</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bukharin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Karampatziakis</surname> <given-names>N</given-names></string-name>, <string-name><surname>He</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AdaLoRA: adaptive budget allocation for parameter-efficient fine-tuning</article-title>. <comment>arXiv:2303.10512. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2303.10512</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Renduchintala</surname> <given-names>A</given-names></string-name>, <string-name><surname>Konuk</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kuchaiev</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Tied-LoRA: enhacing parameter efficiency of lora with weight tying</article-title>. <comment>arXiv:2311.09578. 2023</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2311.09578</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Molchanov</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>YCF</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>KT</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DoRA: weight-decomposed low-rank adaptation</article-title>. <comment>arXiv:2402.09353. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2402.09353</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hyeon-Woo</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ye-Bin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>TH</given-names></string-name></person-group>. <article-title>Fedpara: low-rank hadamard product for communication-efficient federated learning</article-title>. <comment>arXiv:2108.06098. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2108.06098</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name></person-group>. <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. <comment>arXiv:2010.11929. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2010.11929</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Maire</surname> <given-names>M</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hays</surname> <given-names>J</given-names></string-name>, <string-name><surname>Perona</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramanan</surname> <given-names>D</given-names></string-name> <etal>et al</etal></person-group>. <article-title>Microsoft COCO: common objects in context</article-title>. In: <conf-name>Computer Vision&#x2013;ECCV 2014: 13th European Conference</conf-name>; <year>2014 Sep 6&#x2013;12</year>; <publisher-loc>Zurich, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>13</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-10602-1_48</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Plummer</surname> <given-names>BA</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cervantes</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Caicedo</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Hockenmaier</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lazebnik</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>; <year>2015</year>; <publisher-loc>Santiago, Chile</publisher-loc>. p. <fpage>2641</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV.2015.303</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Anne Hendricks</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>O</given-names></string-name>, <string-name><surname>Shechtman</surname> <given-names>E</given-names></string-name>, <string-name><surname>Sivic</surname> <given-names>J</given-names></string-name>, <string-name><surname>Darrell</surname> <given-names>T</given-names></string-name>, <string-name><surname>Russell</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Localizing moments in video with natural language</article-title>. In: <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>; <year>2017</year>; <publisher-loc>Venice, Italy</publisher-loc>. p. <fpage>5803</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV.2017.618</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Goyal</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Khot</surname> <given-names>T</given-names></string-name>, <string-name><surname>Summers-Stay</surname> <given-names>D</given-names></string-name>, <string-name><surname>Batra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Parikh</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Making the V in VQA matter: elevating the role of image understanding in visual question answering</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2017</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>6904</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-018-1116-0</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Video question answering via gradually refined attention over appearance and motion</article-title>. In: <conf-name>Proceedings of the 25th ACM International Conference on Multimedia</conf-name>; <year>2017</year>; <publisher-loc>Mountain View, CA, USA</publisher-loc>. p. <fpage>1645</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3123266.3123427</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>YC</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name>, <string-name><surname>El Kholy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Uniter: universal image-text representation learning</article-title>. In: <conf-name>European Conference on Computer Vision</conf-name>; <year>2020</year>; <publisher-loc>Glasgow, UK</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>104</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-58577-8_7</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>UNIMO: towards unified-modal understanding and generation via cross-modal contrastive learning</article-title>. <comment>arXiv:2012.15409. 2020</comment>. doi:<pub-id pub-id-type="doi">10.18653/v1/2021.acl-long.202</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jia</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>YT</given-names></string-name>, <string-name><surname>Parekh</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pham</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Scaling up visual and vision-language representation learning with noisy text supervision</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2021</year>; <publisher-loc>PMLR</publisher-loc>. p. <fpage>4904</fpage>&#x2013;<lpage>16</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Selvaraju</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gotmare</surname> <given-names>A</given-names></string-name>, <string-name><surname>Joty</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>SCH</given-names></string-name></person-group>. <article-title>Align before fuse: vision and language representation learning with momentum distillation</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>9694</fpage>&#x2013;<lpage>705</lpage>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name> <etal>et al</etal></person-group>. <article-title>Oscar: object-semantics aligned pre-training for vision-language tasks</article-title>. In: <conf-name>Computer Vision&#x2013;ECCV 2020: 16th European Conference</conf-name>; <year>2020 Aug 23&#x2013;28</year>; <publisher-loc>Zurich, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>. p. <fpage>121</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-58577-8_8</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Unifying vision-and-language tasks via text generation</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2021</year>; <publisher-name>PMLR</publisher-name>. p. <fpage>1931</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Seeing out of the box: end-to-end pre-training for vision-language representation learning</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2021</year>. p. <fpage>12976</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01278</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Changpinyo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>N</given-names></string-name>, <string-name><surname>Soricut</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Conceptual 12m: pushing web-scale image-text pre-training to recognize long-tail visual concepts</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2021</year>. p. <fpage>3558</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00356</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VinVL: making visual representations matter in vision-language models</article-title>. <comment>arXiv:2101.00529. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2101.00529</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Scaling up vision-language pre-training for image captioning</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2022</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>17980</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01745</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Berg</surname> <given-names>TL</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Less is more: CLIPBERT for video-and-language learning via sparse sampling</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2021</year>. p. <fpage>7331</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00725</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bain</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nagrani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Frozen in time: a joint video and image encoder for end-to-end retrieval</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <year>2021</year>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>1728</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00175</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Seo</surname> <given-names>PH</given-names></string-name>, <string-name><surname>Nagrani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schmid</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Look before you speak: visually contextualized utterances</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2021</year>. p. <fpage>16877</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01660</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Niebles</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>SC</given-names></string-name></person-group>. <article-title>Align and prompt: video-and-language pre-training with entity prompts</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>; <year>2022</year>. p. <fpage>4953</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.00490</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fu</surname> <given-names>TJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>WY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VIOLET: end-to-end video-language transformers with masked visual-token modeling</article-title>. <comment>arXiv:2111.12681. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2111.12681</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Miech</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sivic</surname> <given-names>J</given-names></string-name>, <string-name><surname>Laptev</surname> <given-names>I</given-names></string-name>, <string-name><surname>Schmid</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Just ask: learning to answer questions from millions of narrated videos</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <year>2021</year>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>1686</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00171</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>KQ</given-names></string-name>, <string-name><surname>Tsutsui</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>All in one: exploring unified video-language pre-training</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2023</year>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>6598</fpage>&#x2013;<lpage>608</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00638</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>W</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Clip4Clip: an empirical study of clip for end to end video clip retrieval and captioning</article-title>. <source>Neurocomputing</source>. <year>2022</year>;<volume>508</volume>:<fpage>293</fpage>&#x2013;<lpage>304</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2022.07.028</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zellers</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hessel</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Park</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MERLOT: multimodal neural script knowledge models</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2021</year>;<volume>34</volume>:<fpage>23634</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>