<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">55560</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.055560</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Efficient User Identity Linkage Based on Aligned Multimodal Features and Temporal Correlation</article-title>
<alt-title alt-title-type="left-running-head">Efficient User Identity Linkage Based on Aligned Multimodal Features and Temporal Correlation</alt-title>
<alt-title alt-title-type="right-running-head">Efficient User Identity Linkage Based on Aligned Multimodal Features and Temporal Correlation</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Gao</surname><given-names>Jiaqi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zheng</surname><given-names>Kangfeng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>kfzheng@bupt.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Xiujuan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Wu</surname><given-names>Chunhua</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Wu</surname><given-names>Bin</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Cyberspace Security, Beijing University of Posts and Telecommunications</institution>, <addr-line>Beijing, 100876</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Faculty of Information Technology, Beijing University of Technology</institution>, <addr-line>Beijing, 100124</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Kangfeng Zheng. Email: <email>kfzheng@bupt.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>15</day><month>10</month><year>2024</year></pub-date>
<volume>81</volume>
<issue>1</issue>
<fpage>251</fpage>
<lpage>270</lpage>
<history>
<date date-type="received">
<day>01</day>
<month>7</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>8</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_55560.pdf"></self-uri>
<abstract>
<p>User identity linkage (UIL) refers to identifying user accounts belonging to the same identity across different social media platforms. Most of the current research is based on text analysis, which fails to fully explore the rich image resources generated by users, and the existing attempts touch on the multimodal domain, but still face the challenge of semantic differences between text and images. Given this, we investigate the UIL task across different social media platforms based on multimodal user-generated contents (UGCs). We innovatively introduce the efficient user identity linkage via aligned multi-modal features and temporal correlation (EUIL) approach. The method first generates captions for user-posted images with the BLIP model, alleviating the problem of missing textual information. Subsequently, we extract aligned text and image features with the CLIP model, which closely aligns the two modalities and significantly reduces the semantic gap. Accordingly, we construct a set of adapter modules to integrate the multimodal features. Furthermore, we design a temporal weight assignment mechanism to incorporate the temporal dimension of user behavior. We evaluate the proposed scheme on the real-world social dataset TWIN, and the results show that our method reaches 86.39% accuracy, which demonstrates the excellence in handling multimodal data, and provides strong algorithmic support for UIL.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>User identity linkage</kwd>
<kwd>multimodal models</kwd>
<kwd>attention mechanism</kwd>
<kwd>temporal correlation</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, with the development of mobile Internet technology and the increase in user demand, there are more and more virtual accounts in cyberspace, and the same user owns multiple accounts in different applications or even on the same platform. UIL refers to identifying user accounts belonging to the same identity across different platforms. UIL is essential for downstream tasks such as information diffusion prediction [<xref ref-type="bibr" rid="ref-1">1</xref>] and cross-platform recommendation [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>Users frequently generate content with concurrent intrinsic relevance across diverse social media platforms, marked by temporal stamps and predominantly manifested as textual or visual content, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. UIL research is bifurcated into approaches leveraging: (i) unimodal data, focusing on text-based elements such as user IDs and posts [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]; and (ii) multimodal data, integrating textual and visual information to holistically profile user identities and behaviors through data complementarity [<xref ref-type="bibr" rid="ref-8">8</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>]. Acknowledging the temporal clustering of contextually related user posts, recent studies [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>] have incorporated temporal dimensions, enhancing the precision of behavioral pattern recognition, rendering user profiles more nuanced, and augmenting cross-platform user identification accuracy and efficacy.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>UIL based on multimodal UGCs</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_55560-fig-1.tif"/>
</fig>
<p>In response to the current state of research, this study focuses on two key points: (1) Exploring an innovative method that utilizes multimodal UGCs. (2) Analyzing the effect of the time factor on user behavior, incorporating the timestamps of UGCs into the model design. For the first point, current UIL methods based on multimodal UGCs generally adopt textual or visual models trained on unimodal data to extract features from heterogeneous data such as text and images, and then directly train feature extraction or modal fusion models based on anchor user pairs to build up an interaction and fusion mechanism between heterogeneous modalities. However, these approaches exist the following problems: (1) Due to the natural semantic gap between text and image modalities, the desired synergistic effect is hard to achieve by integrating multimodal features based on a limited number of anchor pairs across social platforms. (2) The current method needs to train both textual and visual models, and the number of parameters is large, which makes the training inefficient. (3) Different service providers focus on different features, e.g., the content posted on Instagram is mainly images with short text, while the posts posted on Twitter are mainly text supplemented by images, and the text of the posts on Instagram is shorter than Twitter. Coping with the variability of data modality and the imbalance of data volume is another problem to be faced by UIL. For the second point, existing methods usually use a simplified function to summarize the effect of time on user behavior, ignoring the changes that may exist within the period.</p>
<p>To solve the above problems, we propose EUIL. As shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the method includes the BLIP-based text enhancement module, the aligned feature extraction module, the multimodal adapter module, and the temporal correlation-based cross-modal similarity calculation module. Specifically, the BLIP-based text enhancement module first utilizes the powerful multimodal model BLIP [<xref ref-type="bibr" rid="ref-13">13</xref>] to generate captions for images in graphic posts, which not only compensates for the problem of missing textual data but also further explores the deep semantic information of images; then, the alignment feature extraction module unifies the features of texts and images with the pre-trained cross-modal comparative learning model CLIP [<xref ref-type="bibr" rid="ref-14">14</xref>]. Then, the alignment feature extraction module unifies the texts and images with the pre-trained cross-modal contrast learning model CLIP, which can maintain cross-modal semantic consistency while capturing the unique features of the texts and images, effectively bridging the semantic gap between the texts and the images. Furthermore, we design a multimodal adapter module optimized for UIL based on the multi-head self-attention mechanism to generate highly customized user feature representations by fine-tuning limited parameters without changing the underlying model. Finally, the cross-modal similarity computation module based on temporal linkage constructs the inter-user similarity matrix based on the generated multimodal features and dynamically adjusts the importance of the information within each time window in the UIL process by introducing a time factor assignment mechanism. We conduct a series of exhaustive experiments on real datasets to confirm the efficient UIL capability of the proposed method in coping with heterogeneous modal data imbalance.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Illustration of the proposed scheme for UIL</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_55560-fig-2.tif"/>
</fig>
<p>The main contributions of this paper are as follows:
<list list-type="bullet">
<list-item>
<p>Presents an efficient UIL method for multimodal data. Based on the pre-trained multimodal large model CLIP to extract text and image features on social media platforms, the unique features of each text and image are captured while maintaining cross-modal semantic consistency, effectively bridging the semantic gap between text and image, and providing excellent feature initialization for subsequent fine-tuning.</p></list-item>
<list-item>
<p>Innovatively employs the BLIP text generation model to enhance short text on social media platforms. Generating image captions compensates for the problem of insufficient text information and deeply integrates the semantic information of images.</p></list-item>
<list-item>
<p>Presents a multimodal adapter module for UIL tasks. The module effectively fuses text and image features from CLIP based on a multi-head self-attention mechanism. The training of the lightweight multimodal adapter module makes it possible to generate comprehensive feature representations suitable for user identity association without changing the parameters of the pre-trained multimodal model, making the training efficient.</p></list-item>
<list-item>
<p>Innovatively designs a segmented time decay function that distinguishes and meticulously models the differential effects over different periods. This approach not only better fits the complex characteristics of real-world user behavior dynamically changing over time, but also strengthens the model&#x2019;s sensitivity and adaptability to UIL matching under different time windows, thus improving the accuracy of UIL and the practicality of the model.</p></list-item>
<list-item>
<p>Evaluates the proposed method on real social datasets, and the experimental results demonstrate the superiority of the proposed model over state-of-the-art models.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>User Identity Linkage</title>
<p>Existing UIL studies can be categorized into unimodal data-based methods and multimodal data-based methods based on the data modality utilized.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Unimodal Data-Based Methods</title>
<p>Current unimodal UIL methodologies predominantly harness users&#x2019; textual data, encompassing attributes and published content. Based on the fact that people prefer to choose similar usernames across social platforms, EEUPL [<xref ref-type="bibr" rid="ref-3">3</xref>] utilizes usernames and profiles for linkage, assigns similar data to the same bucket via the minHashLSH algorithm, and generates similarity graphs by calculating only the similarity of user pairs within the buckets, which reduces the number of user pairs to be compared. MAUIL [<xref ref-type="bibr" rid="ref-7">7</xref>] categorizes the social media texts into character-, word-, and topic-level attributes, extracting features via unsupervised methodologies tailored to each category. Gao et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] link entities in the texts to a knowledge graph through entity recognition and entity linking techniques, subsequently constructing word similarity matrices and entity similarity matrices using GloVE [<xref ref-type="bibr" rid="ref-15">15</xref>] and TransE [<xref ref-type="bibr" rid="ref-16">16</xref>], respectively. AsyLink [<xref ref-type="bibr" rid="ref-4">4</xref>] first introduces external text location pairs by associating words with locations using the topic modeling technique, reducing the bias caused by sparse link labels. Then it constructs the user-user interaction tensor as the basis for linking and captures the matching patterns in the user interaction tensor with a 3D convolutional neural network. Huang et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] incorporate user attributes, generated content, and check-ins, mitigating local semantic noise via semantic feature extraction and enhancing noise resilience through multi-view graph data augmentation. In addition, graph neural networks are also used for UIL, where the social relationship graph is constructed by considering users as nodes and the relationships between users as edges. To model the influence of the neighbors, MEgo2Vec [<xref ref-type="bibr" rid="ref-17">17</xref>] designs three mechanisms for representing different attributes, distinguishing different neighbors and capturing structural information of the social network, which addresses the problem of error propagation and the presence of noise in social networks. Long et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] propose a degree-aware graph neural network model DegUIL, which effectively solves the challenges brought by tail nodes and super head nodes in the UIL task. By supplementing and correcting the neighborhood information of nodes, DegUIL improves the quality of node representation and thus improves the accuracy of user identity links across social networks.</p>
<p>Unimodal data-based UIL methods rely only on textual modal data for UIL, ignoring the rich semantic information embedded in images, and failing to comprehensively capture and understand users&#x2019; multifaceted behavioral characteristics and personalized preferences.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Multimodal Data-Based Methods</title>
<p>Multimodal data-based methods utilize both textual modal data and visual modal data of users. LinkSocial [<xref ref-type="bibr" rid="ref-8">8</xref>] extracts username features based on bi-grams, crops images to extract faces using OpenFace, an open-source image similarity framework, and uses deep learning to represent the face features on a 128-dimensional unit hypersphere. AHGNet [<xref ref-type="bibr" rid="ref-11">11</xref>] combines several deep learning techniques for extracting user features from multimodal social data, where textual information is captured using (Bidirectional Encoder Representations from Transformers, BERT) [<xref ref-type="bibr" rid="ref-19">19</xref>] and TextCNN [<xref ref-type="bibr" rid="ref-20">20</xref>] fusion models deep semantic features, image information is encoded using ResNet [<xref ref-type="bibr" rid="ref-21">21</xref>] to extract high-dimensional image features, check-in data is represented by constructing spatio-temporal co-occurrence matrix and applying GRU [<xref ref-type="bibr" rid="ref-22">22</xref>] model to form spatio-temporal sequences and social relationships are transformed into feature vectors revealing the structure of the social network using DeepWalk [<xref ref-type="bibr" rid="ref-23">23</xref>] algorithm. UserNet [<xref ref-type="bibr" rid="ref-12">12</xref>] is used to analyze the user&#x0027;s textual information using BiLSTM [<xref ref-type="bibr" rid="ref-24">24</xref>] to encode users&#x2019; text posts to capture contextual information as well as potential emotional and semantic features in the posts, and uses a pre-trained ResNet model to extract the depth features of the images uploaded by users, and then constructs the similarity matrices based on textual features and the similarity matrices based on image features, respectively, and predicts whether the users from the two platforms are the same person by fusing the similarity matrices of the two modalities whether they are the same person or not. AMSA [<xref ref-type="bibr" rid="ref-9">9</xref>] uses the BLIP model to generate textual descriptions for the images and extracts the subject of the text using the Latent Dirichlet Allocation (LDA) [<xref ref-type="bibr" rid="ref-25">25</xref>] model. In addition, AMSA extracts the textual features from the textual collection using BERT and extracts the image features of the user using ConvNeXT [<xref ref-type="bibr" rid="ref-26">26</xref>]. GRU-based model extracts the user&#x2019;s check-in features and fuses them into a user representation vector. MFlink [<xref ref-type="bibr" rid="ref-10">10</xref>] utilizes usernames and user-generated multimodal data to represent the user. Specifically, the method employs the bag-of-words model to extract the username features and extracts the user&#x2019;s posting text and image content information based on the BERT and ResNet models, respectively, and finally integrates the three modalities with the help of graph neural network and attention mechanism to integrate the three modalities. The current research extracts the features of each modal data based on the corresponding unimodal model and then trains the model based on the anchor user data for inter-modal interaction and fusion, however, there is a semantic gap between text and image, and due to the lack of multimodal social data and the difficulty in acquiring the anchor user pairs, it is more difficult to train the feature extraction model and the multimodal fusion model based on text and image features extracted from unimodal models, making the UIL unsatisfactory.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Multimodal Models</title>
<p>Given that visual and linguistic modalities often convey complementary insights, the synergy achieved through joint multimodal representation learning has demonstrated remarkable efficacy across various tasks, including visual question answering, image captioning, and quotation interpretation.</p>
<p>The CLIP (Contrastive Language-Image Pre-training) model is a multimodal pre-training model proposed by OpenAI, designed to bridge the semantic gap between text and images through contrastive learning. The innovation of CLIP lies in its use of unsupervised learning methods, training on large-scale data pairs of image and text scraped from the Internet, thereby learning cross-modal representations. The CLIP model comprises two main components: an image encoder and a text encoder. The image encoder can be any pre-trained convolutional neural network (CNN), such as ResNet, while the text encoder is typically based on the Transformer architecture. These two encoders map input images and text into the same vector space, allowing for a direct comparison of their similarities.</p>
<p>During training, CLIP utilizes a contrastive loss function, which aims to maximize the similarity of positive pairs while minimizing the similarity of negative pairs. Suppose we have a set of images I and corresponding textual descriptions T, where each image <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>i</mml:mi></mml:math></inline-formula> has a corresponding text description <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The loss function can be represented as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x2212;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mfrac><mml:mrow><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>S</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the similarity score between the image <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>i</mml:mi></mml:math></inline-formula> and its corresponding text description <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and &#x1D70F; is the temperature parameter used to scale the similarity scores before applying the softmax function. For each image <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>i</mml:mi></mml:math></inline-formula>, we compute its similarity with all text descriptions, and for each text description <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, we calculate its similarity with all images. Then, we normalize these similarities using the softmax function to obtain probability distribution. The goal of this loss function is to make the similarity of paired image-text descriptions significantly higher compared to non-paired ones. Minimizing this loss function allows the CLIP model to learn meaningful representations in the multimodal space, effectively bridging the semantic gap between text and images.</p>
<p>Building upon CLIP, the BLIP Series Models [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>] represent an advanced multimodal pre-training architecture, incorporating refined training paradigms and architectural enhancements. Beyond joint representation learning, BLIP emphasizes improvements in text generation rooted in visual understanding, thus achieving a harmonious integration of visual analysis and linguistic expression, marking a significant advance toward comprehensive multimodal intelligence. BLIP2 is an upgraded version of BLIP, further enhancing the model&#x0027;s performance and efficiency while retaining its original strengths. BLIP2 achieves higher cross-modal matching accuracy and faster inference speed through improvements in model architecture and optimized training strategies.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Preliminaries</title>
<p>To clearly articulate our research problem and its underlying components, we first introduce the necessary notation and then define the research problem itself.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Notation</title>
<p>We establish a set of symbols that will be used throughout our discussion to ensure clarity and consistency. <xref ref-type="table" rid="table-1">Table 1</xref> summarizes the main notations used in this paper.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of the main notations</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Notation</th>
<th>Explanation</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>The first social media platform.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>The second social media platform.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>The <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>i</mml:mi></mml:math></inline-formula>-th user on <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>The <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>i</mml:mi></mml:math></inline-formula>-th user on <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>The <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>k</mml:mi></mml:math></inline-formula>-th textual post of <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>Image of the <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>g</mml:mi></mml:math></inline-formula>-th post of <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>Text of the <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>g</mml:mi></mml:math></inline-formula>-th post of <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula></td>
<td>Image caption of the <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mi>g</mml:mi></mml:math></inline-formula>-th post of <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>K textual posts by <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>G multimodal posts by <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>.</td>
</tr>
<tr>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula></td>
<td>Expanded posts by <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Problem Formulation</title>
<p>In this work, we aim to address the problem of UIL based on multimodal UGCs across different social media platforms. Without loss of generality, we focus on UIL between <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platforms, and we choose the popular social media Twitter and Instagram as <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platforms for illustration of our scheme. Due to the different service focuses, on Twitter people tend to broadcast news or join topic discussions through text tweets, while on Instagram they tend to post images with short text descriptions to share their daily lives or ideas. To make the model more representative, we specifically assume that users prefer to generate unimodal text posts on <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and multimodal paired image and text posts on <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>Suppose we have a set of training user account pairs <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mrow><mml:mi>&#x1D4B0;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> where each pair consists of two user accounts from different social media platforms (<inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>) are composed of. Also, suppose that for each user account <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> on <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we have his/her K unimodal text posts <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. Similarly, for each user account <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> on <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we collect his/her G multimodal posts <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the image and text of the <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:math></inline-formula> post, respectively. The training user account pairs are labeled by <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi mathvariant="bold-italic">y</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">T</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> stands for the real situation of the first <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>i</mml:mi></mml:math></inline-formula> pair of users <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Specifically, <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> if accounts <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> refer to the same identity in the physical world, and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> otherwise. In a sense, our goal is to use labeled pairs of training user accounts to learn the projection <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>f</mml:mi><mml:mo>&#x003A;</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x00D7;</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Methodology</title>
<p>Facing the task of UIL in multimodal UGCs scenarios, we propose an efficient UIL method based on aligned multimodal features and temporal correlation, including the BLIP-based text enhancement module, the aligned feature extraction module, the multimodal adapter module, and the temporal correlation-based cross-modal similarity calculation module. To address the problem of data modality variability and data volume imbalance in the content posted by users on different social media, the BLIP-based text enhancement module first utilizes the BLIP model to generate a detailed description for the image content as a supplement to the text modality data. Then, to address the semantic gap between image and text, the aligned feature extraction module extracts the representation of text and image under a unified feature space with the CLIP model. Then, we design and train the multimodal adapter module to generate feature representations suitable for UIL. Finally, the temporal correlation-based cross-modal similarity computation module constructs the multimodal user similarity matrix, and we introduce a time factor to reweight the similarity matrix. We input the multimodal similarity matrix into the classification network to classify the input user pairs. The specific implementation of each module is described in detail below.</p>
<sec id="s4_1">
<label>4.1</label>
<title>BLIP-Based Text Enhancement</title>
<p>Diverse social media platforms are characterized by distinct functionalities; for instance, Instagram prioritizes visual content, with user-generated multimodal posts predominantly featuring images accompanied by brief or even absent textual accompaniments. To bolster textual content on such platforms and leverage the copious semantic information embedded in images, we introduce a text augmentation strategy grounded in the BLIP model. This approach taps into the decoder component of a pre-trained BLIP model to generate descriptive captions for individual images within posts. BLIP, a robust bidirectional text-image pre-training model, boasts a decoder capable of translating image substance into coherent natural language narratives. By doing so, it effectively transforms non-verbal image data into textual representations, thereby enabling the interpretation of image content through the lens of linguistic semantics and enriching the overall semantic context.</p>
<p>Specifically, for the <italic>G</italic> multimodal posts <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> for each user <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> on the <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform, a pre-trained BLIP model based on the pre-training BLIP model first used ResNet to process the input image <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> into a high dimensional feature vector. This process captures the spatial and appearance information of the image. Next, the decoder Transformer receives the image feature vector and generates a sequence of word embeddings to form the final caption <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. The caption <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> not only captures the specific visual elements in the image such as objects, scenes, and actions, but also expresses deeper contextual information and implicit relationships. The generated caption <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is paired with the original data <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to get the extended data pair <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which in turn constructs an extended set of posts notated as <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>v</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> contains all the modal information for each post: the image, the original text, and the caption of the image generated by BLIP.</p>
<p>By transforming image information into text, the semantic gap between two different modalities, image and text, can be narrowed, which facilitates the comparison and fusion of data from the two modalities in a unified semantic space, and is conducive to improving the accuracy of cross-modal UIL. Especially on social media platforms, image posts posted by users are often accompanied by short or even no textual descriptions. BLIP-generated captions can effectively supplement this lack of information and provide more valuable information clues for UIL. In addition, image-based caption generation can be viewed as a kind of data augmentation of text, which increases the diversity of data and helps to improve the generalization ability of the subsequent model and the ability to cope with unseen complex situations.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Aligned Feature Extraction</title>
<p>Traditional work tends to use targeted image models and language models to extract image and text features respectively, e.g., extracting 2048-D image feature vectors using the ResNet model and 768-D or 1024-D text features using the BERT model, and then training modal fusion networks or classification networks based on these two features. However, due to the semantic gap between text and images, high quality as well as a high quantity of multimodal data from anchor users is required to achieve the desired training results. Since social data tends to have a lot of noise and anchor users are expensive to acquire, the current results of training fused multimodal features based on unimodal models are not satisfactory.</p>
<p>To address the above problem, we design the alignment feature extraction module to provide a unified feature representation of text and images on social media with the help of a pre-trained cross-modal comparative learning model, CLIP. CLIP employs a comparative learning strategy, the basic idea of which is to maximize the similarity between positive sample pairs (an image and the text that describes it correctly), while minimizing the similarity between negative sample pairs (an image and an irrelevant text description). This approach motivates the model to learn representations that efficiently map images and text into a shared semantic space. The process of extracting features based on the CLLP model can be formalized as <xref ref-type="disp-formula" rid="eqn-2">Eqs. (2)</xref> and <xref ref-type="disp-formula" rid="eqn-3">(3)</xref>:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the image encoder and text encoder of CLIP, respectively. Specifically, for the image <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>v</mml:mi></mml:math></inline-formula> of <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> in the extended set, the image features are obtained using the image encoder <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of CLIP, denoted as <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, with <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> being the <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> output image feature&#x2019;s dimension. For the original text <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>c</mml:mi></mml:math></inline-formula> and image caption <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>d</mml:mi></mml:math></inline-formula> in <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mrow><mml:mover><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, the text features are obtained using CLIP&#x0027;s text encoder <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, respectively, denoted as <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, with <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> being the dimension of the <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> output text feature. Similarly, for text <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>t</mml:mi></mml:math></inline-formula> in the set of <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the text encoder of CLIP is also utilized to obtain its text features <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, which are more comparable between multimodal features <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> than those extracted by the unimodal model. CLIP can capture the unique features of each of the text and image while maintaining cross-modal semantic consistency, so that content expressing the same concept or topic will have similar representations even in different modalities, effectively bridging the semantic gap between text and image and providing high-quality input features for subsequent UIL.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Multimodal Adapter</title>
<p>The contents posted on different social media platforms by users are often different, and for the UIL task, some of the posts are extremely similar in content, which possesses a positive impact on the UIL effect, while some of the posts are equivalent to noise, which harms the UIL effect. Considering the different degrees of contribution of each post to end-UIL, we design a lightweight Adapter module to personalize and optimize the features extracted from the CLIP model to further enhance the performance of these features in the UIL task. The Adapter module is typically a pre-trained model that is supplemented by one or more layers of a small network structure, allowing for the targeting of the UIL without changing the parameters of the original model. Adapter modules typically add one or more layers of small network structures to a pre-trained model, allowing fine-tuning of the model&#x0027;s performance for a specific task without changing the original model parameters. In the multimodal UIL scenario, we design the multimodal Adapter module as a four-terminal, multi-headed self-attention module, with each end as shown in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mover><mml:mi>Z</mml:mi><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi></mml:math></inline-formula> stands for Multi-Head Self-Attention Operation, <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> stands for Multilayer Vector Machine Operation, <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>Z</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the text feature of user <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are the image feature, text feature, and image caption feature of user <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mrow><mml:mover><mml:mi>Z</mml:mi><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> which correspond to the outputs of the features in <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>Z</mml:mi></mml:math></inline-formula> after passing through the Adapter module, respectively.</p>
<p>Specifically, we set up the Adapter module for the text end of the <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform, the text end of the <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform, the image end, and the caption end, respectively. First, the features of each end are passed into an MLP network, as shown in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>Z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>Z</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The MLP consists of two fully connected layers, where <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the weight matrices of the first and second layers, respectively, and <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the bias terms of the corresponding layers. The <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the activation function. The MLP layer performs more complex nonlinear transformations on the multimodal features extracted by CLIP to fully explore and combine the high-level abstraction relationships among the features.</p>
<p>The feature <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> output from the MLP network is then passed to the Multihead Self-Attention module, and the Multihead Self-Attention is calculated as shown in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mover><mml:mi>Z</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>H</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>O</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mi>h</mml:mi></mml:math></inline-formula> denotes the number of heads, <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the output of the <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>i</mml:mi></mml:math></inline-formula>-<italic>th</italic> head, and <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>O</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the output transformation matrix. The output of each head <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be calculated by <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:math></inline-formula> is the attention computation function, <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are the query, key, and value transformation matrices of the <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mi>i</mml:mi></mml:math></inline-formula><italic>-th</italic> head, <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the dimension of the key vector <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:math></inline-formula> normalizes the similarity, the weight of each key vector is computed, and then multiply the weights by the value vectors, and finally perform a weighted summation to get the Attention output.</p>
<p>The Adapter module utilizes a multi-head self-attention mechanism to assign dynamic weights to the features of each modality of each post so that those post features that are more representative and distinguishable in determining the identity of a user will receive higher attention and weighting. The multi-head self-attention mechanism processes the input features in parallel from multiple different perspectives, and each head outputs a weighted subset of feature vectors, and then stitches together the results from the individual heads to ultimately generate a new representation of the features that reflect the importance of each post.</p>
<p>The multimodal Adapter module makes full use of the powerful cross-modal features provided by CLIP and fine-tunes the differentiation of the importance of image features based on the actual application scenarios, which improves the accuracy of UIL, and does not require fine-tuning of the large multimodal model during training, but only needs to train the lightweight Adapter module, which improves the performance of UIL.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Temporal Correlation-Based Cross-Modal Similarity</title>
<p>In response to the behavioral patterns of users posting content on different social media platforms, it is found that users tend to post content with some kind of intrinsic relevance within the same period. Based on this observation, we construct a multimodal similarity computation framework incorporating a time factor for measuring users&#x0027; content consistency across platforms.</p>
<p>First, we construct a user similarity matrix based on text posts on the <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform and multimodal posts containing text, images, and captions on the <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform as shown in <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the cosine similarity between the image content of <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and the text content of <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the cosine similarity between the text content of <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the cosine similarity between <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>&#x0027;s caption content and <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>&#x0027;s text content. <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> measures the content consistency of users between the two platforms.</p>
<p>However, considerations based solely on content similarity are not sufficient to fully capture the temporal nature of user behavior. Given that content posted by users within the same period is more relevant, we introduce a time factor <italic>r</italic> to adjust the similarity weights. This time factor <italic>r</italic> is defined as shown in <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mn>2</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mi>x</mml:mi></mml:math></inline-formula> represents the publishing time difference between the two posts, <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> are the set time difference threshold parameters, and <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> control the magnitude of the change in the decay function. The image of the time decay function is shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, when the two posts are closer to each other in terms of publishing time, the larger the time factor value is, and the higher the corresponding similarity confidence is. Specifically, when the publishing time difference <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mi>x</mml:mi></mml:math></inline-formula> between two posts is small enough, i.e., <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mn>0</mml:mn><mml:mo>&#x003C;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we assign a reward value greater than 1 to the similarity between the posts, and as the time difference increases, the time factor gradually decays, but because the time difference is still small at this time, the time factor decays slowly. When the time difference <italic>x</italic> between two posts is large, i.e., <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mi>x</mml:mi><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we assign a penalty value less than 1 to the similarity between the posts, and as the time difference increases, the time factor gradually decays, and it decays rapidly, converging to 0. When the time difference is within the fuzzy period, i.e., <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we retain the original content similarity, i.e., we set the time factor to 1.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Image of the time decay function. The horizontal coordinate is the time interval and the vertical coordinate is the time factor value</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_55560-fig-3.tif"/>
</fig>
<p>According to <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>, we calculate the time factor matrix <italic>R</italic> between <italic>K</italic> textual posts and <italic>G</italic> multimodal posts, and finally, we combine the time factor-adjusted cross-modal similarity defined as a list of content similarities weighted by the time factor as shown in <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> as follows:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="negativethinmathspace" /><mml:mo>,</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>S</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:msup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msup><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mo>&#x2299;</mml:mo></mml:math></inline-formula> denotes multiplication by elements. The final prediction is shown in <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>o</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>A</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:msup><mml:mi>m</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mi>A</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> performs mean pooling on the final similarity matrix to reduce the dimensionality, thus forming a vector that reflects the overall similarity between user <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and user <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Finally, this dimensionally reduced user similarity vector is input to the fully connected layer network for classification, with <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denoting the weight of the classification layer, <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denoting the bias of the classification layer, and the output of the fully connected layer transformed into a probability distribution by the <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:math></inline-formula> function. After optimized training, the proposed network can effectively predict whether the user pairs from two platforms are the same person.</p>
<p>The cross-modal similarity calculation method based on the time factor not only effectively addresses the limitations of single-modal similarity calculations but also successfully leverages temporal information to enhance the accuracy and effectiveness of the UIL task. By incorporating the time factor into the similarity calculation, this method provides a more nuanced understanding of the relationship between different modalities, particularly when dealing with unaligned image-text pairs. Algorithm 1 outlines the complete process of the EUIL algorithm.</p>
<fig id="fig-5">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_55560-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experiments</title>
<p>To demonstrate the effectiveness of the proposed method, we conducted experiments on real-world datasets to prove the effectiveness of the proposed method. All experiments were conducted on servers with Intel (R) Xeon (R) Gold 6330 CPUs and three Nvidia A800 GPUs.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Setup</title>
<p>Dataset. We validate our approach on the TWIN dataset [<xref ref-type="bibr" rid="ref-12">12</xref>] collected by Chen et al. The TWIN dataset collects information about users on two popular heterogeneous social media platforms, Twitter and Instagram. After filtering out some low-quality data, 5765 user pairs were obtained on Twitter and Instagram, along with 1,729,500 UGCs and corresponding timestamps.</p>
<p>Training settings. We divided the user account pairs into three parts: 80% for training, 10% for validation, and 10% for testing. These are considered as positive samples, while negative samples are randomly generated with the same number of positive and negative samples. We converted the UIL into a binary classification task, using accuracy as the evaluation metric. For optimization, we used the Adam optimizer with a learning rate of 0.0001.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Model Comparison</title>
<p>Due to the limited research done on the issue of linking user identities in UGCs, we compare our scheme with the following baseline:
<list list-type="bullet">
<list-item>
<p>WSF-GBDT: This baseline is derived from the method in [<xref ref-type="bibr" rid="ref-28">28</xref>], which introduced four writing-style features, including lexical, syntactic, structural, and content-specific features, to characterize users and employed the support vector machine [<xref ref-type="bibr" rid="ref-29">29</xref>] for the UIL.</p></list-item>
<list-item>
<p>WHOLE: This baseline also characterizes each user account&#x2019;s textual/visual modality by considering all UGCs as a whole. WHOLE employs BiLSTM to derive the textual representation, conducts the text-text similarity and the text-image similarity separately by MLP, and obtains the final user similarity by fusing the text-text similarity and the text-image similarity.</p></list-item>
<list-item>
<p>UserNet: This baseline utilizes two bidirectional long and short-term memory networks to encode text posts from users of the <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platforms, extracts the depth features of the images uploaded by users of the <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platforms using a pre-trained ResNet model, and then constructs the similarity matrices based on the text features and the similarity matrix based on image features, and makes introduce time factor to weight the similarity. The similarity matrices of the two modalities are fused through the attention mechanism to predict whether the users from the two platforms have the same identity.</p></list-item>
</list></p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> shows the performance comparison of the different methods. Based on this table, we obtain the following observations: (1) Our proposed method and UserNet are significantly better than WSF-GBDT. this may be attributed to the fact that the former two methods integrate multimodal user representations, not just writing style features. (2) Our proposed method outperforms WHOLE, demonstrating the advantages of fine-grained similarity modeling in the context of user identity correlation, and can further incorporate temporal correlation. (3) Our proposed method significantly outperforms the optimal baseline UserNet. This may be attributed to the fact that UserNet extracts text features using the text-specific unimodal model BiLSTM and image features using the visual unimodal model ResNet, respectively, and computes the similarity matrices directly on top of the two unimodal features, ignoring the semantic gaps between the text and image. semantic gap between text and image. In contrast, our proposed method, which uses the text encoder and image encoder of the multimodal model CLIP to extract features, allows the features of different modalities to be in the same embedding space that can be compared and is more capable of capturing the similarity relationship between image and text. To visually demonstrate the effectiveness of the multimodal model in extracting aligned features on social media, we use a similarity matrix to show the similarity of the image-text features extracted by CLIP, where the lighter the color of the block of elements in the matrix, the higher the image similarity. Where image-text pairs come from heterogeneous multimodal posts made by users on social media, there is often a potential correlation between images and text. As shown in the <xref ref-type="fig" rid="fig-4">Fig. 4</xref> below, the features extracted by the CLIP model better capture this correlation.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance comparison among different models in terms of accuracy</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td>WHOLE</td>
<td>0.6836</td>
</tr>
<tr>
<td>WSF-GBDT</td>
<td>0.7552</td>
</tr>
<tr>
<td>UserNet</td>
<td>0.8369</td>
</tr>
<tr>
<td>EUIL (ours)</td>
<td><bold>0.8639</bold></td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Similarity matrix of multimodal posts based on features extracted from CLIP model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_55560-fig-4.tif"/>
</fig>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Ablation Experiments</title>
<p>To gain a deeper understanding of our proposed model, we compare our proposed method EUIL with several derived methods. Specifically, we explore the effects of modal diversity, text enhancement, multimodal Adapter, and time factor on UIL performance.
<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Only the text modal of the <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform is considered and only the original text posted by the user is taken into account, ignoring the image modal data of the <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform with the text of the generated image description.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Only the image modal of the <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform is considered, ignoring the text modal data of the <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform, which includes the original text posted by the user with the generated caption.</p></list-item>
<list-item>
<p>EUIL<sub><italic>text&#x0026;img</italic></sub>: Only the raw data posted by the users of the <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform is considered, including the raw text data posted by the users along with the image data, ignoring the captions generated for the images.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Only the image modality of the <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform and the captions generated for the images were considered, ignoring the raw text posted by the users of the <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: We disable the multimodal Adapter module of EUIL to directly construct a similarity matrix based on the features extracted from the pre-trained model CLIP to predict whether the users have the same identity.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: We disable the time factor of EUIL and unweight the time factor on the similarity when calculating the similarity between posts.</p></list-item>
</list></p>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>On Modal</title>
<p>To explore the contribution of each modality to UIL, we compared the results between EUIL, <inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which are shown in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Results on the effect of different modal combinations on the accuracy</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>K &#x003D; G &#x003D; 150</th>
<th>K &#x003D; G &#x003D; 100</th>
<th>K &#x003D; G &#x003D; 50</th>
</tr>
</thead>
<tbody>
<tr>
<td>EUIL</td>
<td>0.8815</td>
<td>0.8671</td>
<td>0.8191</td>
</tr>
<tr>
<td><inline-formula id="ieqn-192"><mml:math id="mml-ieqn-192"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.7654</td>
<td>0.7446</td>
<td>0.7134</td>
</tr>
<tr>
<td><inline-formula id="ieqn-193"><mml:math id="mml-ieqn-193"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.6549</td>
<td>0.5717</td>
<td>0.5709</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The primary observation that the accuracy of the EUIL model is higher than that of <inline-formula id="ieqn-188"><mml:math id="mml-ieqn-188"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> may be rooted in the fact that the <inline-formula id="ieqn-189"><mml:math id="mml-ieqn-189"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> model is limited to analyzing only the textual information of the <inline-formula id="ieqn-190"><mml:math id="mml-ieqn-190"><mml:msub><mml:mrow><mml:mi>&#x1D4AA;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> platform, whereas the EUIL model not only integrates the textual information but also incorporates the rich resources of image data. Image data, as a complement, can fill in the potential contextual gaps and detail omissions in textual descriptions, and this fine-grained information is often difficult to fully convey by pure textual representations. Second, the observation that the accuracy of EUIL is higher than that of the <inline-formula id="ieqn-191"><mml:math id="mml-ieqn-191"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> model emphasizes the central role of textual data in decoding the user&#x0027;s intention and contextual framework. User-generated text is rich in direct emotional expressions and specific points of view, elements that are essential for fine-grained characterization of users&#x0027; personalities and pinpointing their needs. Taken together, these analyses point to the conclusion that the fusion of multiple data modalities (e.g., text and images) is decisive for building a more comprehensive user profile, which significantly enhances the performance and accuracy of user-matching models. This conclusion emphasizes the significance of cross-modal data integration in enhancing the effectiveness of UIL systems.</p>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>On Image Caption</title>
<p>To validate the contribution of the BLIP-based text enhancement module to UIL, we compare the results of EUIL, <inline-formula id="ieqn-194"><mml:math id="mml-ieqn-194"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-195"><mml:math id="mml-ieqn-195"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-196"><mml:math id="mml-ieqn-196"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and the results are shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Results on the effect of captions on the accuracy</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>K &#x003D; G &#x003D; 150</th>
<th>K &#x003D; G &#x003D; 100</th>
<th>K &#x003D; G &#x003D; 50</th>
</tr>
</thead>
<tbody>
<tr>
<td>EUIL</td>
<td><bold>0.8815</bold></td>
<td><bold>0.8671</bold></td>
<td><bold>0.8191</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-197"><mml:math id="mml-ieqn-197"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.8647</td>
<td>0.8471</td>
<td>0.7862</td>
</tr>
<tr>
<td><inline-formula id="ieqn-198"><mml:math id="mml-ieqn-198"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.5124</td>
<td>0.5116</td>
<td>0.4604</td>
</tr>
<tr>
<td><inline-formula id="ieqn-199"><mml:math id="mml-ieqn-199"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.6549</td>
<td>0.5717</td>
<td>0.5709</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Preliminary observations show that the accuracy of EUIL model outperforms <inline-formula id="ieqn-200"><mml:math id="mml-ieqn-200"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, while the performance of <inline-formula id="ieqn-201"><mml:math id="mml-ieqn-201"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is significantly higher than that of <inline-formula id="ieqn-202"><mml:math id="mml-ieqn-202"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This finding strongly proves the validity of the textual descriptions of the images and indicates that the textual descriptions of the images deeply mine and enrich the semantic content behind the image, which in turn enhances the expressiveness of user features. In addition, the superior performance of <inline-formula id="ieqn-203"><mml:math id="mml-ieqn-203"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>g</mml:mi><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> further highlights the strategy of expanding textual data by integrating textual descriptions of images, which is especially crucial to compensate for the data limitations of social media platforms that rely on a single textual modality. In summary, our proposed BLIP-Enhanced Textual Enrichment module significantly enhances the accuracy and efficacy of user matching by generating textual descriptive information of images, validating its effectiveness in enhancing UIL.</p>
</sec>
<sec id="s5_3_3">
<label>5.3.3</label>
<title>On Multimodal Adapter</title>
<p>To validate the contribution of the multimodal Adapter module to UIL, we compared the results of EUIL and <inline-formula id="ieqn-204"><mml:math id="mml-ieqn-204"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which are shown in <xref ref-type="table" rid="table-5">Table 5</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Results on the effect of multimodal adapter on accuracy</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>K &#x003D; G &#x003D; 150</th>
<th>K &#x003D; G &#x003D; 100</th>
<th>K &#x003D; G &#x003D; 50</th>
</tr>
</thead>
<tbody>
<tr>
<td>EUIL</td>
<td><bold>0.8815</bold></td>
<td><bold>0.8671</bold></td>
<td><bold>0.8191</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-205"><mml:math id="mml-ieqn-205"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.5532</td>
<td>0.5468</td>
<td>0.5460</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Observations show that the EUIL model exhibits a significant advantage in accuracy over its variant <inline-formula id="ieqn-206"><mml:math id="mml-ieqn-206"><mml:msub><mml:mrow><mml:mtext>&#xA0;EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This performance difference may be rooted in the simplifying assumption that <inline-formula id="ieqn-207"><mml:math id="mml-ieqn-207"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is adopted in processing the characteristics of K/G user posts, i.e., it defaults to the same importance of the information contained in all the posts, which fails to adequately take into account the intrinsic differences and noise specific to user-generated content (UGC) in different social media platforms. On the contrary, the EUIL model explicitly identifies the information noise embedded in the K/G samples through deep insights and is highly sensitive to the uneven contribution of each post to the user matching task, which plays a central role in the innovative multimodal adapter mechanism introduced by EUIL, which dynamically and adaptively assigns differentiated weights to each user&#x0027;s K/G posts. This strategy not only optimizes the adjustment based on the relevance and quality of each post but also effectively utilizes multimodal information to enhance feature representation. Therefore, EUIL not only refines the learning of each post&#x0027;s features but also significantly improves the accuracy and robustness of UIL through this series of fine-tuning, verifying its effectiveness and superiority.</p>
</sec>
<sec id="s5_3_4">
<label>5.3.4</label>
<title>On Time Factor</title>
<p>To validate the contribution of introducing the time factor to UIL, we compared the results of EUIL with <inline-formula id="ieqn-208"><mml:math id="mml-ieqn-208"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which are shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Results on the effect of time factor on the accuracy</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>K &#x003D; G &#x003D; 150</th>
<th>K &#x003D; G &#x003D; 100</th>
<th>K &#x003D; G &#x003D; 50</th>
</tr>
</thead>
<tbody>
<tr>
<td>EUIL</td>
<td><bold>0.8815</bold></td>
<td><bold>0.8671</bold></td>
<td><bold>0.8191</bold></td>
</tr>
<tr>
<td><inline-formula id="ieqn-209"><mml:math id="mml-ieqn-209"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.8639</td>
<td>0.8311</td>
<td>0.7958</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Observations show that the accuracy of the EUIL model outperforms that of the <inline-formula id="ieqn-210"><mml:math id="mml-ieqn-210"><mml:msub><mml:mrow><mml:mtext>EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> model, a difference that may stem from the fundamental difference in their approaches to user posts. The <inline-formula id="ieqn-211"><mml:math id="mml-ieqn-211"><mml:msub><mml:mrow><mml:mtext>&#xA0;EUIL</mml:mtext></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> model evaluates only the direct similarity in the content of users&#x0027; posts, ignoring the effect of the time dimension. In contrast, the EUIL model incorporates a temporal dimension into the consideration of content relevance, i.e., it assumes that there is an intrinsic relevance of the content posted by users in a similar period. By integrating the time information of posts to construct a temporal relevance model, EUIL effectively enhances the granularity of post-similarity modeling, which in turn optimizes the effectiveness of user matching. This finding highlights the importance of integrating time-series features with content features to enhance UIL accuracy in user behavior analysis.</p>
</sec>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Discussion</title>
<sec id="s6_1">
<label>6.1</label>
<title>Potential Improvements</title>
<p>Our approach leverages aligned multimodal features and temporal correlation to enhance the performance of the UIL task. By utilizing aligned multimodal features and incorporating temporal information, we have achieved significant improvements. However, despite these advancements, there remain some limitations and areas for potential improvement:</p>
<p>(1) Integration of Multi-Dimensional User Features: EUIL primarily focuses on studying the contribution of multi-modal user-generated content to user alignment. However, different social media platforms emphasize various types of content, which poses significant challenges to the proposed methods, potentially affecting their generalization and robustness. Beyond user-generated content, other user features on social media platforms, such as user profiles and social network structures, may provide crucial clues for enhancing the accuracy of UIL. To build more universally applicable models, future work should consider incorporating a wider range of user features into the model.</p>
<p>(2) Construction of Datasets with Multi-Dimensional User Features: Training and evaluating models using more diverse datasets ensures their generality and robustness. However, due to user privacy protection, publicly available datasets that meet the requirements are limited. Therefore, constructing datasets that include comprehensive features is necessary for future work, including collecting data from multiple social media platforms, and covering information from diverse user groups, thereby enhancing the representativeness and breadth of research. In this way, we can more comprehensively evaluate the model&#x0027;s performance in real-world scenarios, ensuring its adaptability to real-world environments and improving its recognition capabilities in complex situations.</p>
<p>(3) The dynamic nature of social media user behavior and preferences poses significant challenges for UIL. Capturing long-term trends and short-term changes is crucial for maintaining the accuracy and relevance of these systems. In future work, we can pursue research along the following three directions: (1) Longitudinal Tracking Mechanism: We plan to develop a mechanism capable of tracking user behavior over extended periods. This will involve collecting and analyzing historical data to identify patterns and trends in user activities. By doing so, the system can adapt to long-term changes in user behavior. (2) Dynamic Learning Algorithms: We aim to design algorithms that can continuously update user models based on incoming data. This dynamic learning capability will enable the system to quickly adapt to short-term changes in user behavior, such as shifts in interests or sudden changes in activity patterns. (3) Feedback Loop for Model Refinement: We propose implementing a feedback loop to refine user models based on real-time feedback. This could involve engaging users with the system, such as through corrections or verifications of linked identities, to improve the accuracy of the models. We believe that these enhancements will significantly improve the robustness and adaptability of our UIL system. They will enable the system to maintain its performance even as user behavior evolves.</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Privacy Security</title>
<p>The performance of EUIL on the TWIN dataset indicates that UIL relying solely on UGC across Twitter and Instagram platforms is feasible. Unlike user profiles, which can be disclosed or contain false information, UGC often represents genuine user behavior and is publicly available on social platforms, which makes UGC-based UIL techniques susceptible to malicious exploitation for consolidating sensitive personal information across social media platforms. Therefore, it is necessary to take measures to ensure privacy security. Research on data privacy protection in other fields is relatively advanced [<xref ref-type="bibr" rid="ref-30">30</xref>], but research on user social data privacy protection is relatively insufficient. There are specific UIL adversarial attack methods targeting user attributes, such as the DeLink [<xref ref-type="bibr" rid="ref-31">31</xref>] framework, which draws from adversarial text generation ideas to help users modify their social media usernames, thus defending against UIL. Additionally, there are adversarial attack methods targeting the structure of social networks, such as the TOAK [<xref ref-type="bibr" rid="ref-32">32</xref>] strategy, which leverages the topological structure information of networks by carefully perturbing network structures to reduce the matching accuracy of UIL models. However, there are currently adversarial methods specifically targeting UGC. To prevent UGC from being maliciously utilized, our next steps will explore how to generate adversarial UGC by introducing imperceptible perturbations without disrupting the intended expression of the user, thereby reducing the effectiveness of UGC-based UIL strategies and further protecting user data privacy.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Conclusion</title>
<p>In this paper, we propose an efficient UIL scheme based on aligned multimodal features and temporal correlation to address the challenges of multimodal UIL in social media platforms. The method utilizes the BLIP model for image-generated captions, which compensates for the lack of textual information in social media platforms and significantly improves the expressive power of image modality in UIL tasks. Moreover, the method makes full use of the pre-trained multimodal large model CLIP to realize the representation of data in both text and image modalities under a unified feature space, achieving an effective fusion of cross-modal features. The subsequently designed multimodal adapter module, by its lightweight and flexible features, can extract a comprehensive feature representation that is highly correlated and suitable for user identity association from the multimodal features extracted by CLIP by training with a small number of additional parameters. In addition, we introduce a time factor assignment mechanism to construct a user similarity prediction model that can dynamically reflect the temporal dynamics of user interest evolution and social behavior, thus enhancing the timeliness and accuracy of the UIL method. Experiments demonstrate that this method performs well on large-scale real-world social media datasets, not only overcoming the problem of data modal differences and imbalance but also surpassing the existing state-of-the-art methods in terms of accuracy and efficiency.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec><title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization and methodology, Jiaqi Gao; software, Jiaqi Gao and Bin Wu; validation, Xiujuan Wang, Kangfeng Zheng and Chunhua Wu; writing&#x2014;original draft preparation, Jiaqi Gao; writing&#x2014;review and editing, Xiujuan Wang and Chunhua Wu; supervision, Kangfeng Zheng. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the corresponding author, Kangfeng Zheng, upon reasonable request.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Trajcevski</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Zhong</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<chapter-title>Information diffusion prediction via recurrent cascades convolution</chapter-title>,&#x201D; <article-title>presented at the 2019 IEEE 35th Int. Conf. Data Eng. (ICDE)</article-title>, <publisher-loc>Macao, China</publisher-loc>, <year>Apr. 8&#x2013;11, 2019</year>, pp. <fpage>770</fpage>&#x2013;<lpage>781</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T. H.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gao</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>CROSS: Cross-platform recommendation for social E-commerce</article-title>,&#x201D; <conf-name>presented at the Proc. 42nd Int. ACM SIGIR (Special Interest Group Inf. Retr.) Conf. Res. Develop. Inf. Retr.</conf-name>, <publisher-loc>Paris, France</publisher-loc>, <year>Jul. 21&#x2013;25, 2019</year>, pp. <fpage>515</fpage>&#x2013;<lpage>524</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>EEUPL: Towards effective and efficient user profile linkage across multiple social platforms</article-title>,&#x201D; <source>World Wide Web</source>, vol. <volume>24</volume>, no. <issue>5</issue>, pp. <fpage>1731</fpage>&#x2013;<lpage>1748</lpage>, <year>Jun. 2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s11280-021-00882-7</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shen</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Cheng</surname></string-name></person-group>, &#x201C;<article-title>AsyLink: User identity linkage from text to geo-location via sparse labeled data</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>515</volume>, no. <issue>6</issue>, pp. <fpage>174</fpage>&#x2013;<lpage>184</lpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2022.10.027</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Xing</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wu</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>A Semantic-enhancement-based social network user-alignment algorithm</article-title>,&#x201D; <source>Entropy</source>, vol. <volume>25</volume>, no. <issue>1</issue>, <year>Jan. 2023</year>, Art. no. 172. doi: <pub-id pub-id-type="doi">10.3390/e25010172</pub-id>; <pub-id pub-id-type="pmid">36673313</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Gao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Shao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shen</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Cheng</surname></string-name></person-group>, &#x201C;<article-title>User identity linkage across social networks with the enhancement of knowledge graph and time decay function</article-title>,&#x201D; <source>Entropy</source>, vol. <volume>24</volume>, no. <issue>11</issue>, <year>Nov. 2022</year>, Art. no. 1603. doi: <pub-id pub-id-type="doi">10.3390/e24111603</pub-id>; <pub-id pub-id-type="pmid">36359693</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>MAUIL: Multilevel attribute embedding for semisupervised user identity linkage</article-title>,&#x201D; <source>Inf. Sci.</source>, vol. <volume>593</volume>, pp. <fpage>527</fpage>&#x2013;<lpage>545</lpage>, <year>May 2022</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ins.2022.02.023</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>V.</given-names> <surname>Sharma</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Dyreson</surname></string-name></person-group>, &#x201C;<chapter-title>LINKSOCIAL: Linking user profiles across multiple social media platforms</chapter-title>,&#x201D; <article-title>presented at the 2018 IEEE Int. Conf. Big Knowl. (ICBK)</article-title>, <publisher-loc>Singapore</publisher-loc>, <year>Nov. 17&#x2013;18, 2018</year>, pp. <fpage>260</fpage>&#x2013;<lpage>267</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Gou</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Xiong</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Cui</surname></string-name></person-group>, &#x201C;<chapter-title>The potential utility of image descriptions: User identity linkage across social networks based on MultiModal self-attention fusion</chapter-title>,&#x201D; <article-title>presented at the 2023 IEEE Int. Perform., Comput., Commun. Conf. (IPCCC)</article-title>, <publisher-loc>Anaheim, CA, USA</publisher-loc>, <year>Nov. 17&#x2013;19, 2023</year>, pp. <fpage>265</fpage>&#x2013;<lpage>273</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>MFLink: User identity linkage across online social networks via multimodal fusion and adversarial learning</article-title>,&#x201D; <source>IEEE Trans. Emerg. Top. Comput. Intell.</source>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>Mar. 2024</year>. doi: <pub-id pub-id-type="doi">10.1109/TETCI.2024.3440057</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Feng</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Nie</surname></string-name></person-group>, &#x201C;<chapter-title>Adversarial-enhanced hybrid graph network for user identity linkage</chapter-title>,&#x201D; <article-title>presented at the Proc. 44th Int. ACM SIGIR Conf. Res. Develop. Inf. Retr.</article-title>, <year>Jul. 11&#x2013;15, 2021</year>, pp. <fpage>1084</fpage>&#x2013;<lpage>1093</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Cui</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Gan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Cheng</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Nie</surname></string-name></person-group>, &#x201C;<article-title>User identity linkage across social media via attentive time-aware user modeling</article-title>,&#x201D; <source>IEEE Trans. Multimedia</source>, vol. <volume>23</volume>, pp. <fpage>3957</fpage>&#x2013;<lpage>3967</lpage>, <year>Nov. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TMM.2020.3034540</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Xiong</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Hoi</surname></string-name></person-group>, &#x201C;<chapter-title> BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation</chapter-title>,&#x201D; <article-title>presented at the Proc. 39th Int. Conf. Mach. Learn. (PMLR)</article-title>, <publisher-loc>Baltimore, MD, USA</publisher-loc>, <year>Jul. 17&#x2013;23, 2022</year>, pp. <fpage>12888</fpage>&#x2013;<lpage>12900</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Radford</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>Learning transferable visual models from natural language supervision</chapter-title>,&#x201D; <article-title>presented at the Proc. 38th Int. Conf. Mach. Learn. (ICML)</article-title>, <year>Jul. 18&#x2013;24, 2021</year>, pp. <fpage>8748</fpage>&#x2013;<lpage>8763</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Pennington</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Socher</surname></string-name>, and <string-name><given-names>C. D.</given-names> <surname>Manning</surname></string-name></person-group>, &#x201C;<chapter-title>GloVe: Global vectors for word representation</chapter-title>,&#x201D; <article-title>presented at the Proc. 2014 Conf. Empir. Methods Nat. Lang. Process. (EMNLP)</article-title>, <publisher-loc>Doha, Qatar</publisher-loc>, <year>Oct. 25&#x2013;29, 2014</year>, pp. <fpage>1532</fpage>&#x2013;<lpage>1543</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Bordes</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Usunier</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Garcia-Duran</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Weston</surname></string-name>, and <string-name><given-names>O.</given-names> <surname>Yakhnenko</surname></string-name></person-group>, &#x201C;<chapter-title>Translating embeddings for modeling multi-relational data</chapter-title>,&#x201D; <article-title>presented at the Proc. 26th Int. Conf. Neural Inf. Proc. Syst.</article-title>, <publisher-loc>Lake Tahoe, NV, USA</publisher-loc>, <year>Dec. 05&#x2013;10, 2013</year>, pp. <fpage>2787</fpage>&#x2013;<lpage>2795</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>MEgo2Vec: Embedding matched ego networks for UIL across social networks</chapter-title>,&#x201D; <article-title>presented at the Proc. 27th ACM Int. Conf. Inf. Knowl. Manag.</article-title>, <publisher-loc>Torino, Italy</publisher-loc>, <year>Oct. 22&#x2013;26, 2018</year>, pp. <fpage>327</fpage>&#x2013;<lpage>336</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Long</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title> DegUIL: Degree-aware graph neural networks for long-tailed user identity linkage</chapter-title>,&#x201D; <article-title>presented at the Proc. Eur. Conf. Mach. Learn. Princ. Pract. Knowl. Discov. Databases</article-title>, <publisher-loc>Turin, Italy</publisher-loc>, <year>Sep. 18&#x2013;22, 2023</year>, pp. <fpage>122</fpage>&#x2013;<lpage>138</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Devlin</surname></string-name>, <string-name><given-names>M. W.</given-names> <surname>Chang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Toutanova</surname></string-name></person-group>, &#x201C;<article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>,&#x201D; <year>2018</year>, <comment><italic>arXiv:1810.04805</italic></comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Convolutional neural network for sentence classification</article-title>,&#x201D; <year>2015</year>, <comment><italic>arXiv:1408.5882</italic></comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<chapter-title>Deep residual learning for image recognition</chapter-title>,&#x201D; <article-title>presented at the Proc. 2016 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</article-title>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>Jun. 27&#x2013;30, 2016</year>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Chung</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Empirical evaluation of gated recurrent neural networks on sequence modeling</article-title>,&#x201D; <year>2014</year>, <comment><italic>arXiv:1412.3555</italic></comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Perozzi</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Al-Rfou</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Skiena</surname></string-name></person-group>, &#x201C;<chapter-title>DeepWalk: Online learning of social representations</chapter-title>,&#x201D; <article-title>presented at the Proc. 20th ACM SIGKDD Int. Conf. Knowl. Discov. Data Min.</article-title>, <publisher-loc>NY, USA</publisher-loc>, <year>Aug. 24&#x2013;27, 2014</year>, pp. <fpage>701</fpage>&#x2013;<lpage>710</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Hochreiter</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Schmidhuber</surname></string-name></person-group>, &#x201C;<article-title>Long short-term memory</article-title>,&#x201D; <source>Neural Comput.</source>, vol. <volume>9</volume>, no. <issue>8</issue>, pp. <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>, <year>Nov. 1997</year>. doi: <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id>; <pub-id pub-id-type="pmid">9377276</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D. M.</given-names> <surname>Blei</surname></string-name>, <string-name><given-names>A. Y.</given-names> <surname>Ng</surname></string-name>, and <string-name><given-names>M. I.</given-names> <surname>Jordan</surname></string-name></person-group>, &#x201C;<article-title>Latent dirichlet allocation</article-title>,&#x201D; <source>J. Mach. Learn. Res.</source>, vol. <volume>3</volume>, pp. <fpage>993</fpage>&#x2013;<lpage>1022</lpage>, <year>Mar. 2003</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Mao</surname></string-name>, <string-name><given-names>C. Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Feichtenhofer</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Darrell</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Xie</surname></string-name></person-group>, &#x201C;<chapter-title>A convnet for the 2020s</chapter-title>,&#x201D; <article-title>presented at the Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</article-title>, <publisher-loc>New Orleans, LA, USA</publisher-loc>, <year>Jun. 18&#x2013;24, 2022</year>, pp. <fpage>11976</fpage>&#x2013;<lpage>11986</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models</chapter-title>,&#x201D; <article-title>presented at the Proc. Int. Conf. Mach. Learn. (ICML)</article-title>, <publisher-loc>Honolulu, HI, USA</publisher-loc>, <year>Jul. 23&#x2013;29, 2023</year>, pp. <fpage>19730</fpage>&#x2013;<lpage>19742</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>A framework for authorship identification of online messages: Writing style features and classification techniques</article-title>,&#x201D; <source>J. Am. Soc. Inf. Sci. Technol.</source>, vol. <volume>57</volume>, no. <issue>3</issue>, pp. <fpage>378</fpage>&#x2013;<lpage>393</lpage>, <year>Dec. 2005</year>. doi: <pub-id pub-id-type="doi">10.1002/asi.20316</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Cortes</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Vapnik</surname></string-name></person-group>, &#x201C;<article-title>Support-vector networks</article-title>,&#x201D; <source>Mach. Learn.</source>, vol. <volume>20</volume>, no. <issue>3</issue>, pp. <fpage>273</fpage>&#x2013;<lpage>297</lpage>, <year>Sep. 1995</year>. doi: <pub-id pub-id-type="doi">10.1007/BF00994018</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Buccafurri</surname></string-name>, <string-name><given-names>V. De</given-names> <surname>Angelis</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Lazzaro</surname></string-name></person-group>, &#x201C;<article-title>MQTT-A: A broker-bridging P2P architecture to achieve anonymity in MQTT</article-title>,&#x201D; <source>IEEE Internet Things J.</source>, vol. <volume>10</volume>, no. <issue>17</issue>, pp. <fpage>15443</fpage>&#x2013;<lpage>15463</lpage>, <year>Sep. 2023</year>. doi: <pub-id pub-id-type="doi">10.1109/JIOT.2023.3264019</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>DeLink: An adversarial framework for defending against cross-site user identity linkage</article-title>,&#x201D; <source>ACM Trans. Web</source>, vol. <volume>18</volume>, no. <issue>2</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>34</lpage>, <year>Mar. 2024</year>. doi: <pub-id pub-id-type="doi">10.1145/3643828</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Shao</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>TOAK: A topology-oriented attack strategy for degrading user identity linkage in cross-network learning</chapter-title>,&#x201D; <article-title>presented at the Proc. 32nd ACM Int. Conf. Inf. Knowl. Manag. (CIKM)</article-title>, <publisher-loc>Birmingham, UK</publisher-loc>, <year>Oct. 21&#x2013;25, 2023</year>, pp. <fpage>2208</fpage>&#x2013;<lpage>2218</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>