<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">60395</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.060395</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>MMCSD: Multi-Modal Knowledge Graph Completion Based on Super-Resolution and Detailed Description Generation</article-title>
<alt-title alt-title-type="left-running-head">MMCSD: Multi-Modal Knowledge Graph Completion Based on Super-Resolution and Detailed Description Generation</alt-title>
<alt-title alt-title-type="right-running-head">MMCSD: Multi-Modal Knowledge Graph Completion Based on Super-Resolution and Detailed Description Generation</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wang</surname><given-names>Huansha</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>whs123@mail.ustc.edu.cn</email></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Huang</surname><given-names>Ruiyang</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>gisexpert@163.com</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Qinrang</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Shaomei</given-names></name></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Jianpeng</given-names></name></contrib>
<aff id="aff-1"><institution>National Digital Switching System Engineering &#x0026; Technological R&#x0026;D Center, Information Engineering University</institution>, <addr-line>Zhengzhou, 450001</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Huansha Wang. Email: <email>whs123@mail.ustc.edu.cn</email>; Ruiyang Huang. Email: <email>gisexpert@163.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>26</day><month>03</month><year>2025</year>
</pub-date>
<volume>83</volume>
<issue>1</issue>
<fpage>761</fpage>
<lpage>783</lpage>
<history>
<date date-type="received">
<day>31</day>
<month>10</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>1</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_60395.pdf"></self-uri>
<abstract>
<p>Multi-modal knowledge graph completion (MMKGC) aims to complete missing entities or relations in multi-modal knowledge graphs, thereby discovering more previously unknown triples. Due to the continuous growth of data and knowledge and the limitations of data sources, the visual knowledge within the knowledge graphs is generally of low quality, and some entities suffer from the issue of missing visual modality. Nevertheless, previous studies of MMKGC have primarily focused on how to facilitate modality interaction and fusion while neglecting the problems of low modality quality and modality missing. In this case, mainstream MMKGC models only use pre-trained visual encoders to extract features and transfer the semantic information to the joint embeddings through modal fusion, which inevitably suffers from problems such as error propagation and increased uncertainty. To address these problems, we propose a Multi-modal knowledge graph Completion model based on Super-resolution and Detailed Description Generation (MMCSD). Specifically, we leverage a pre-trained residual network to enhance the resolution and improve the quality of the visual modality. Moreover, we design multi-level visual semantic extraction and entity description generation, thereby further extracting entity semantics from structural triples and visual images. Meanwhile, we train a variational multi-modal auto-encoder and utilize a pre-trained multi-modal language model to complement the missing visual features. We conducted experiments on FB15K-237 and DB13K, and the results showed that MMCSD can effectively perform MMKGC and achieve state-of-the-art performance.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multi-modal knowledge graph</kwd>
<kwd>knowledge graph completion</kwd>
<kwd>multi-modal fusion</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<award-id>BHQ090003000X03</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Multi-modal knowledge graph (MMKG) refers to the large-scale graph-structured knowledge base that utilizes semantic knowledge networks to describe multi-modal entities in the form of triples like <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">d</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Numerous studies have shown that multi-modal data can introduce more supervised information and external knowledge into deep neural networks. Thus, MMKGs have become one of the hottest research fields of knowledge graphs, with important applications in multi-modal knowledge question answering system, retrieval-augmented generation, and so on. Due to the continuous growth of data and knowledge, as well as the quality limitations of online encyclopedias (which are the main knowledge source of knowledge graphs), the existing mainstream knowledge graphs inevitably have problems such as insufficient knowledge and incomplete content. A specific manifestation is the lack of entities or relations in the triples like <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">d</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">r</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">o</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>?</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> or <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold-italic">h</mml:mi><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">d</mml:mi></mml:mrow><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>?</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">a</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">l</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mi mathvariant="bold-italic">n</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">i</mml:mi><mml:mi mathvariant="bold-italic">t</mml:mi><mml:mi mathvariant="bold-italic">y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The absence of entities or relations in the triples will lead to a break in the knowledge chain, affecting the accuracy of reasoning and searching for knowledge. In order to correctly fill in the missing content in the triples and thus improve the richness of the knowledge, more and more researchers are focusing on Multi-modal Knowledge Graph Completion (MMKGC). MMKGC task aims to discover the unknown triples in the MMKG and complete the entities and relations that should be contained in incomplete triples, which is one of the keys to the construction and extension of the multi-modal knowledge graph.</p>
<p>Mainstream MMKGC models are accomplished based on representation learning, which encodes attributes and relations of each modality into low-dimensional embedding. Subsequently, various modal fusion methods are employed to obtain multi-modal embeddings for both entities and relations, and joint embeddings are formed for all candidate triples, and then decoded to obtain the correct results. Xie et al. [<xref ref-type="bibr" rid="ref-1">1</xref>] used the pre-trained AlexNet to extract visual features from entity description images and sought consistency of the vector space by translation between multi-modal embeddings. Mousselly Sergieh et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] adopted Imagine and DeViSE to fuse multi-modal information. Wang et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed a specific multi-modal auto-encoder module to extract multi-modal joint embeddings directly. Some researchers [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>] also focus on mitigating inter-modal interference and improving the quality of multi-modal embedding by modifying the sampling strategy, refining the modal embedding, or introducing the large language model.</p>
<p>However, existing models generally ignore the poor quality and lack of multi-modal information present in mainstream datasets as well as in reality MMKGs. As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, not every multi-modal entity has visual attributes that highly correspond to its semantics, and this is mainly reflected in the following aspects: (1) Not all entities have multi-modal information, and for entities without visual information, mainstream models typically use random, mean, or zero tensors as visual embeddings, which has less positive effect on the modal interaction and modeling. (2) The semantic relationship between multi-modal information and the entity itself is poor. Unlike the image-text matching task, the text modality information of entities in the MMKGC task may only be the entity name (some datasets have a few sentences of entity description), making it difficult to align the text modality and visual modality well. (3) Multi-modal information itself has quality issues. Due to the fact that mainstream MMKGC datasets are built on semi-structured data from open-source encyclopedias, many entity description images themselves are not clear enough, and fail to effectively describe entities. In this case, previous MMKGC models only use pre-trained visual models to extract visual features, and transfer the information to the joint embeddings through modal fusion, which inevitably have problems such as error propagation and increased uncertainty. Zhang et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] have recognized the negative impact of missing modalities and have proposed the use of generative adversarial networks (GAN) to supplement the missing features. However, training generative adversarial networks consumes a lot of additional computational resources and is disconnected from the original triples. Moreover, it does not solve the problem of poor visual features due to the low quality of the described images.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Three different quality levels of multi-modal entities and their impact on downstream tasks</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-1.tif"/>
</fig>
<p>To address the above issues, we propose MMCSD, a Multi-modal knowledge graph Completion model based on Super-resolution and Detailed Description Generation. The core idea is to optimize the visual information in the joint embedding by improving the quality of the visual embedding as well as completing the missing features, thus facilitating modal interaction. Specifically, we utilize the pre-trained residual network [<xref ref-type="bibr" rid="ref-8">8</xref>] to enhance the resolution and improve the quality of the visual modality. Moreover, we design multi-level visual semantic extraction and entity description generation, thereby further extracting entity semantics from structural triples and visual images. Meanwhile, we train a variational multi-modal auto-encoder and utilize the pre-trained multi-modal language model to complement the missing visual features. We conduct experiments to verify the effectiveness of the model on FB15K-237 [<xref ref-type="bibr" rid="ref-9">9</xref>] and DB13K. The results show that MMCSD can effectively perform MMKGC and achieve state-of-the-art performance.</p>
<p>The contributions of this paper are as follows:
<list list-type="order">
<list-item>
<p>To address the problem of poor visual embedding due to the low quality of descriptive images and primitive way of visual feature extraction, we propose to utilize a pre-trained residual network to super-resolve the original images and employ the multi-level visual semantic extraction mechanism to extract the deep semantics hidden in the visual modality.</p></list-item>
<list-item>
<p>To address the problem of modality missing, we introduce the modality imagination mechanism, i.e., by training the variational multi-modal auto-encoder to simulate the missing visual embedding. Meanwhile, we propose to adopt the multi-modal large model to generate features semantically similar to the triples as a complement.</p></list-item>
<list-item>
<p>We structure the knowledge of DB13K in the multi-modal knowledge graph MMKB<xref ref-type="fn" rid="fn1"><sup>1</sup></xref><fn id="fn1">
<label>1</label>
<p><ext-link ext-link-type="uri" xlink:href="https://github.com/mniepert/mmkb">https://github.com/mniepert/mmkb</ext-link> (accessed on 11 May 2018).</p>
</fn>
(Multi-modal Knowledge Bases) [<xref ref-type="bibr" rid="ref-10">10</xref>] to construct a multi-modal knowledge completion dataset for model evaluation. The dataset will be available at Github.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Knowledge Graph Completion</title>
<p>With the proposal of knowledge graph, researchers have long recognized the importance of knowledge completeness for knowledge graph applications, and thus knowledge graph completion has been one of the most critical related tasks. Traditional knowledge completion tasks mainly rely on manual effort. Although this method has high quality, with the continuous increase in the scale of knowledge graphs, manual completion is no longer applicable to existing knowledge graphs. Representation learning technology excels at extracting strong features and learning optimal representations of objects, thereby simplifying downstream task steps and improving their effectiveness. Due to the ability to extract strong features from complex data forms, representation learning has important applications in the knowledge graph, mainly including translational-distance based models and neural network based models.</p>
<p>Knowledge completion based on translational-distance involves two steps. First, it projects entities onto a low dimensional vector space and treats the relations between entities as translations of entity vectors. Then, it models entities and relations separately through iterative training. The earliest translational-distance based model is TransE [<xref ref-type="bibr" rid="ref-11">11</xref>]. It demonstrates its excellent modeling ability in sparse knowledge graph work and inspires many similar works based on translational-distance in the future. However, the assumptions are too simple, which makes it difficult to model complex relations of one-to-many and many-to-many better. TransH [<xref ref-type="bibr" rid="ref-12">12</xref>] makes the same entity or relation have different vector representations in different triples by introducing hyperplanes instead of relationship vectors. RotatE [<xref ref-type="bibr" rid="ref-13">13</xref>], on the other hand, treats relations as rotation angles between entity vectors, thus modeling entities and relations in more complex space. Translational-distance based models have the advantages of simplicity and a low number of parameters, but accordingly, their effectiveness generally fails to match that of neural network based models. Currently, translational-distance based models are usually used as the basis for other extended models, the entity and relation embeddings pre-trained by TransE are employed as the original features for training.</p>
<p>Moreover, with deep learning technology demonstrating its powerful feature extraction capabilities in fields such as computer vision and natural language processing, knowledge graph completion based on neural networks has become mainstream. ConvE [<xref ref-type="bibr" rid="ref-14">14</xref>] uses convolutional neural networks to aggregate entities and relations, and then calculate corresponding similarity scores. ConvKB [<xref ref-type="bibr" rid="ref-15">15</xref>] associates a unique feature embedding for different entities and relations to improve the transfer and representation ability of the model. With the continuous development of graph convolutional networks (GCN), researchers have attempted to utilize the advantages of GCNs in learning node and edge representations in graph-structured data to better learn entity-relation features of knowledge graphs. R-GCN [<xref ref-type="bibr" rid="ref-16">16</xref>] introduces the relation matrix in GCNs as a mapping transform when entities aggregate neighborhood features. KBAT [<xref ref-type="bibr" rid="ref-17">17</xref>] utilizes graph attention network as the encoder to integrate the information of multi-hop neighbors of entities and then uses ConvKB to decode entity representations. In recent years, a large number of knowledge completion models have emerged that adopt Transformer [<xref ref-type="bibr" rid="ref-18">18</xref>] or pre-trained models to model triples, and have achieved good results. Neural network based models can effectively extract hidden potential features from the knowledge graph with high accuracy, high inference scalability and efficiency. However, neural networks rely on a large amount of training data, which is a data-driven endeavor that usually performs poorly when applied to knowledge graphs with sparse data. In addition, these models suffer from common shortcomings of neural networks such as low interpretation and too many parameters.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Multi-Modal Knowledge Graph Completion</title>
<p>With the continuous development of multi-modal technology, researchers have gradually realized the positive effects of multi-modal data such as images and videos on improving the completeness and universality of knowledge graphs. Many multi-modal knowledge graphs based on large-scale image training sets, single-modal knowledge graphs, Wikipedia, and other data, such as ImgPedia [<xref ref-type="bibr" rid="ref-19">19</xref>], VisualSem [<xref ref-type="bibr" rid="ref-20">20</xref>], GAKG [<xref ref-type="bibr" rid="ref-21">21</xref>], have emerged in large numbers. At the same time, multi-modal knowledge completion tasks have also become one of the important subtasks of knowledge completion.</p>
<p>IKRL [<xref ref-type="bibr" rid="ref-1">1</xref>] for the first time introduces visual information to the knowledge graph completion task by using a pre-trained AlexNet to extract visual features from entity description images and seeks consistency in the vector space by transforming between multi-modal embeddings. MKBE [<xref ref-type="bibr" rid="ref-22">22</xref>] uses different encoders to process numerical and textual data to generate corresponding embeddings in addition to description images, respectively, thus introducing more supervisory information. TransAE [<xref ref-type="bibr" rid="ref-3">3</xref>] processes pre-generated visual and text features through multi-modal auto-encoders to obtain multi-modal joint embeddings as entity representations. These methods are ineffective in extracting and representing multi-modal information, thus affecting the final embedding quality.</p>
<p>Since the modal interactions between the MMKGC tasks are still unclear, researchers have tried several mechanisms to facilitate them. VBKGC [<xref ref-type="bibr" rid="ref-23">23</xref>] adopts a transformer-based multi-modal pretraining model to simultaneously process multi-modal image-text pairs to encode entities. MANS [<xref ref-type="bibr" rid="ref-4">4</xref>] proposes modality aware negative sampling strategy to generate modality-level negative samples where the descriptive images are not correlated with the entity. MoSE [<xref ref-type="bibr" rid="ref-5">5</xref>] focuses on the difference of modality importance. It learns modality-split relation embeddings for each modality, which alleviates the modality interference. MACO [<xref ref-type="bibr" rid="ref-7">7</xref>] leverages the generative adversarial framework and trains a pair of generator-discriminators to generate missing modality features, preventing degradation of embedding quality due to missing multi-modal information. KoPA [<xref ref-type="bibr" rid="ref-6">6</xref>] integrates pre-trained structural embeddings with large language models (LLMs), and achieves structural-aware reasoning in the LLMs.</p>
<p>However, previous research has not focused on further exploration and mining for visual modality processing, and it is still common to use only pre-trained encoders to obtain visual embeddings. This simple visual feature extraction approach does not take into account the difference in relevance between images and entities, as well as the quality of the images themselves. Given the generally poor quality of multi-modal information in datasets and real-world MMKGs, the use of this approach will have a negative impact on the modeling of visual embeddings and joint embeddings. Moreover, most existing approaches do not consider the problem of missing modalities, and for entities with missing visual modality, mainstream models typically use random, mean, or zero tensors as their alternative visual embeddings, which do not contribute positively to modal interaction and fusion. Although there has been work that recognizes detrimental effects of modality missing and proposes to generate auxiliary embeddings using generative adversarial networks, this approach still has limitations. Training generative adversarial networks requires significant additional computational resources, and the auxiliary embeddings are divorced from the knowledge of the entities themselves and thus lack interpretation.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Preliminary</title>
<p>A multi-modal knowledge graph (MMKG) can be denoted as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>E</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>A</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>E</mml:mi></mml:math></inline-formula> is the set of entities, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>R</mml:mi></mml:math></inline-formula> is the set of relations, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>V</mml:mi></mml:math></inline-formula> is the set of visual knowledge, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>T</mml:mi></mml:math></inline-formula> is the set of triples, and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>A</mml:mi></mml:math></inline-formula> is the set of entity attributes (note that in some datasets, entity attributes are also provided in triple form). A triple can be represented as (<inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>E</mml:mi></mml:math></inline-formula> are head entity and tail entity, and <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>r</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>R</mml:mi></mml:math></inline-formula> is relation.</p>
<p>Multi-modal knowledge graph completion (MMKGC) models represent entities and relations as embeddings, and then use a score function <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to assess the likelihood of triples. For the evaluation query <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo>?</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> or <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo stretchy="false">(</mml:mo><mml:mo>?</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the model will sort all candidate entities and output a ranked list of preferences.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Method</title>
<p>To solve the problem of modality missing and low quality visual embedding in MMKGC, we propose MMCSD, a multi-modal knowledge graph completion model based on super-resolution and detailed description generation. It follows the established framework of previous MMKGC models, which involves first extracting embeddings for different modalities separately, then generating multi-modal embeddings through modal fusion, and finally decoding the concatenated features of all candidate triples.</p>
<p>On this basis, we use a pre-trained residual network to improve the resolution of the visual modalities and prevent the degradation of the quality of the visual embedding due to the low clarity of the original image. Compared to the traditional approach of extracting features using only visual encoders, we employ a multi-level visual semantic extraction mechanism to comprehensively capture the semantic information within the visual modality by combining the outputs of multiple mature models. We also feed the triples into the large language model, aiming to obtain knowledge from a natural language perspective beyond structural features. For the missing modality problem, we train a variational multi-modal auto-encoder to simulate the missing visual embeddings and adopt a multi-modal large language model to supplement the original visual features.</p>
<p>In this chapter, we elaborate on the specific architecture and composition of MMCSD. <xref ref-type="sec" rid="s4_1">Sections 4.1</xref>&#x2013;<xref ref-type="sec" rid="s4_3">4.3</xref> detail the generation of embeddings for three different modalities respectively, with a particular emphasis on the visual embeddings. Following that, <xref ref-type="sec" rid="s4_4">Sections 4.4</xref> and <xref ref-type="sec" rid="s4_5">4.5</xref> describe the modality fusion and decoding mechanisms. The missing modality completion mechanism we propose is explained in <xref ref-type="sec" rid="s4_6">Section 4.6</xref>.</p>
<p>The entire multi-modal embedding generation process of MMCSD is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>The multi-modal embedding generation process of MMCSD. As detailed in Chapter 4, we adopt VGG (Visual Geometry Group) and CLIP (Contrastive Language-Image Pre-Training) as visual encoder, SpGAT as structural encoder, and BLIP2 (Bootstrapping Language-Image Pre-training with frozen unimodal models) as attribute encoder</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-2.tif"/>
</fig>
<sec id="s4_1">
<label>4.1</label>
<title>Neighborhood Structural Embedding</title>
<p>Most mainstream knowledge completion models use graph convolutional networks (GCN) [<xref ref-type="bibr" rid="ref-24">24</xref>] or graph attention networks (GAT) [<xref ref-type="bibr" rid="ref-25">25</xref>] to extract entity structural information from knowledge graph triples.</p>
<p>GAT exhibits excellent feature extraction capabilities for graph-structured data, but it also entails relatively high computational complexity. In MMKGC, the relevant edges of most entities are sparse, meaning that the encoder only needs to compute the weights between a relatively small number of nodes for each entity. Therefore, in order to effectively obtain structural embeddings while conserving computational resources, we utilize Sparse GAT (SpGAT) [<xref ref-type="bibr" rid="ref-17">17</xref>] to calculate the attention weights of each relevant triple and iteratively obtain multi-hop knowledge.</p>
<p>Specifically, for entity <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, in order to obtain its structural embedding <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, SpGAT firstly initializes its related triple embeddings <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> through embedding concatenation and linear operations:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>vectors <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote original embeddings of entities <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and relation <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of triple <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, respectively. <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> denotes the linear transformation matrix.</p>
<p>Similar to vanilla GAT, SpGAT then quantifies and calculates the importance <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and attention weight <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of each triple for entity structural embedding:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>k</mml:mi><mml:mi>y</mml:mi><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> refer to a weight matrix, <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>L</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>k</mml:mi><mml:mi>y</mml:mi><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi></mml:math></inline-formula> represent the corresponding functions.</p>
<p>Subsequently, the model sums up the triple embeddings based on the calculated attention weights <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to obtain final structural entity embeddings <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the neighborhood of entity <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the set of relations connecting entities <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<p>We use multi-head attention to get more comprehensive and complete knowledge about the neighborhood. SpGAT employs averaging calculation results of <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>M</mml:mi></mml:math></inline-formula> attention mechanisms to get final embedding vectors for entities:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>M</mml:mi></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mtext>&#xA0;&#xA0;&#xA0;&#xA0;&#xA0;&#xA0;</mml:mtext><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Visual Embedding</title>
<p>Visual modality, as the most unique modality in MMKGs, contains a large amount of semantics related to entities or triples. The mainstream models, which only use pre-trained visual models (such as VGG or ResNet) to encode visual information, have significant limitations in the situation of low visual modality quality in the MMKG. To address this issue, we incorporate an additional step of adopting a residual network to perform super-resolution on the visual modality prior to using a pre-trained visual encoder to extract original visual embeddings. Following this, we leverage a variety of well-established techniques from the fields of computer vision and multi-modal learning to extract visual semantics at multiple levels. This approach enhances the semantic similarity between structural embeddings and visual embeddings, thereby facilitating subsequent modal interaction and fusion.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Original Visual Embedding</title>
<p>We use a pre-trained visual model (PVM), e.g., VGG-16 (Visual Geometry Group 16-layer network), to encode the described images of the entities, and then adopt the final layer output of it as the original visual features.</p>
<p>Afterward, the original visual features are input into a feed-forward neural network to obtain the original visual embeddings with the same dimension as the entity structural features:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mi>M</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the original visual embeddings of <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msup><mml:mi>W</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are learnable weight and the bias matrix of corresponding feed-forward neural network, <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>P</mml:mi><mml:mi>V</mml:mi><mml:mi>M</mml:mi></mml:math></inline-formula> means pre-trained visual model, <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the visual image of <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. When a single entity has multiple descriptive images, we apply mean pooling to aggregate visual features.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Super-Resolution</title>
<p>The resolution of an image greatly affects the accuracy of its semantic analysis. Most mainstream MMKGs are constructed based on open-source crowd-sourcing semi-structured encyclopedias, where the sources of visual modality are extensive and complex. It leads to certain visual modalities within the knowledge graph having low clarity and resolution. Some MMKGs are even unable to provide original images and only present pre-trained visual features. Taking the dataset FB15K-237 as an example, MMKGC models that do not focus on original feature extraction typically only utilize its visual features pre-extracted by VGG. This over-reliance on general visual encoders is based on an unrealistic assumption that all visual modalities and entity ontologies have high semantic similarity, and it ignores the issue of low modality quality, greatly affecting the effectiveness of modal fusion.</p>
<p>We first consider addressing this issue from the perspective of improving the quality of visual modalities. Most of the original images in the datasets are only 30&#x2013;50 KB in size. Therefore, we adopt super-resolution technology. This helps enhance the resolution of the images in the MMKG and optimize the subsequent semantic extraction and modeling effects. Super-resolution technology aims to preserve the semantics of the original image while increasing its resolution to enhance its clarity. EDSR [<xref ref-type="bibr" rid="ref-8">8</xref>] is a renowned work in the field of super-resolution, which enhances performance by removing redundant structures in traditional residual network based super-resolution methods. Although it was an early work, it was a state-of-the-art work at that time, and did not consume too much computing resources. Therefore, considering all factors, we use EDSR as the super-resolution tool in the model architecture.</p>
<p>Considering that the entity description images in the MMKB, which is the data source for FB15K-237 and DB13K, usually come from three different sources (Google, Bing, and Yahoo) and are often repeated. In order to save computational resources while maintaining the comparability of the final experimental results, we do not simply perform super-resolution on all images (FB15K-237 has approximately 270,000 entities, each with 0&#x2013;15 descriptive images). Instead, we pre-use VGG-16 to extract semantic features from each descriptive image and compare them with the visual features of the entities provided in the dataset. Due to the small size of the images within the graph, extracting semantic features does not require extensive computational resources or time. Then, we select the image with the highest similarity for processing and super-resolution.</p>
<p>Specifically, we adopt EDSR to perform super-resolution on the target image <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which increases the length and width by four times each, in order to obtain a clear image <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> that preserves semantic information. Afterwards, we use the multi-modal pre-trained model CLIP to extract the features of <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, serving as a supplement to <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. The primary reason for choosing CLIP is to maintain vector space consistency with the embeddings generated in the subsequent missing modality generation mechanism, and CLIP also has better visual semantic extraction capability than VGG-16:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mi>m</mml:mi><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Multi-Level Visual Semantic Extraction</title>
<p>Most previous models simply used pre-trained visual encoders to extract visual features, which only allowed for coarse-grained semantic extraction of semantically rich visual modality, resulting in insufficient semantic similarity between visual embeddings and other embeddings. To address this issue, inspired by Monkey [<xref ref-type="bibr" rid="ref-26">26</xref>], we propose a multi-level visual semantic extraction mechanism. It aims to create rich and high-quality image descriptions by effectively mixing outputs from various generators.</p>
<p>Similar to Monkey, we use several advanced technologies or systems to combine: BLIP2 [<xref ref-type="bibr" rid="ref-27">27</xref>], which understands image semantics and generates descriptive text, PPOCR [<xref ref-type="bibr" rid="ref-28">28</xref>] extracts possible text information from images based on its powerful optical character recognition capability, GRiT [<xref ref-type="bibr" rid="ref-29">29</xref>] identifies targets and their positions in the image and performs detailed image text matching, FastSAM [<xref ref-type="bibr" rid="ref-30">30</xref>] segments images based on semantics, and Qwen1.5 [<xref ref-type="bibr" rid="ref-31">31</xref>], utilizing its powerful contextual understanding and generation capabilities.</p>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the pipeline for multi-level visual semantic extraction: (a) We first extract the name of the entity from the datasets as the core of the entire mechanism, hoping that the content generated will not deviate too much. (b) Subsequently we use BLIP2 to generate the global descriptions of entity visual information so as to analyze and describe the specific content of the image from an overall perspective. (c) Next, we adopt GRiT to identify specific areas and object coordinates in the image and generate detailed descriptions, PPOCR to extract possible textual information, FastSAM to segment objects. Then BLIP2 is adopted to generate descriptions of them. (d) A filter based on BLIP2 for image matching is used to delete low-confidence objects and areas recognized by GRiT and FastSAM. As for optical characters, we directly use the confidence output by PPOCR to screen them. (e) Finally, detailed information including global description, text extraction, and objects with spatial coordinates are fed into Qwen1.5 for fine-tuning, enabling it to generate semantically rich descriptive text. For the final generated summary description text, we use BLIP2 to extract its semantic features <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> as the other auxiliary to the original visual features.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The pipeline for the multi-level visual semantic extraction</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-3.tif"/>
</fig>
<p>When constructing the prompt for inputting Qwen1.5, we imitated GRiT<xref ref-type="fn" rid="fn2"><sup>2</sup></xref><fn id="fn2">
<label>2</label>
<p><ext-link ext-link-type="uri" xlink:href="https://github.com/JialianW/GRiT">https://github.com/JialianW/GRiT</ext-link> (accessed on 01 December 2022).</p>
</fn>  by inputting captions into ChatGPT to generate image scene descriptions, and additionally informing it of the entity name and the task of generating entity descriptions, therefore achieving good generation results.</p>
<p>By aggregating the unique features of these systems and models, the multi-level visual semantic extraction mechanism captures the details in the visual modality and generates concise, accurate, and contextual descriptive text. Thus supplementing the entity images that were originally semantically sparse, and preventing poor multi-modal fusion caused by low modality quality.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Triple Semantic Embedding</title>
<p>As the main component of the MMKGC datasets, triple data contains the most important attributes and relations of entities. Previous models only extract adjacency relationships between entity nodes using graph neural networks. Based on these relationships, they generate structural embeddings for entity nodes and relationship edges. However, triples still embody a large amount of semantic information and the graph neural network cannot fully understand them.</p>
<p>Some studies have similarly recognized this issue and proposed using language models to extract the semantics of entity names as a supplement. Yet, the semantic information contained solely in entity names is not sufficiently rich and can easily lead to semantic confusion with entities that have similar names. All triples related to an entity contain all the knowledge associated with that entity in MMKGs. Therefore, we attempt to leverage large language models to extract all the entity-related semantics contained within them.</p>
<p>In detail, we input the name and all triples of each entity into the Qwen1.5 and prompt it to generate a small text description. The triple is highly correlated with the entity, and there must be a head or tail entity in the triple that is consistent with the target entity. Therefore, using the LLM to analyze triples for generating entity description yields better results. Small-scale testing revealed that when the number of triples related to an entity is three or more, the text description generated by the LLM retains the semantics of the triples effectively without introducing excessive noise.</p>
<p>A specific example of attribute embedding generation is provided in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. We also use BLIP2 to extract features of text descriptions generated based on triples, treating them as alternative attribute embeddings <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and serving as an external supplement to structural and visual embeddings:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>a</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>B</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi><mml:mn>2</mml:mn><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mi>w</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mn>1.5</mml:mn><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Modal Fusion</title>
<p>Multi-modal fusion aims to integrate information from different modalities into a unified representation. Following the previous works, we project the structural, visual, and attribute embeddings of entities onto the same vector space based on the translational-distance function used in TransE. Afterwards, each uni-modal embedding is further input into a single-layer neural network for weighted fusion to get original multi-modal embedding <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, promoting them to also meet the translational-distance strategy.</p>
<p>For a specific triple <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the uni-modal translational-distance function can be denoted as:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Due to the general difficulty in finding appropriate visual images to describe the abstract relations within triples, simple feature concatenation or weighted summation is generally not used alone in modal fusion in MMKGC that focus on entity-entity relationships. Therefore, after obtaining the initial multi-modal embeddings through weighted summation, we train the network to ensure that the structural, visual, attribute, and joint multi-modal embeddings all meet the translational-distance strategy, thereby promoting spatial consistency among the embeddings:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mi mathvariant="normal">&#x0026;</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>a</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>We additionally use the Margin Ranking Loss function, hoping that the model&#x2019;s ranking score for positive samples is higher than that for the negative samples generated by randomly modifying the head or tail entities:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula> is a margin hyper-parameter, <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>S</mml:mi></mml:math></inline-formula> is the set of positive triples, <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the set of negative triples, and <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> refer to the multi-modal joint embedding obtained through weighted sum.</p>
<p>The final loss function can be expressed as:
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Decoder</title>
<p>We use ConvE to decode the triple embeddings obtained by concatenating multi-modal joint embeddings of entities and relations. For the triple <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the decoding process can be denoted as:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>V</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>g</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>&#x03C9;</mml:mi></mml:math></inline-formula> represents the convolutional filter, <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:math></inline-formula> is the convolution operator, <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>f</mml:mi></mml:math></inline-formula> are the activation function and <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents a linear transformation matrix used to compute the final score of the triple. Inspired by KBAT, we adopt the approach that fuses the final trained entity embeddings with the original embeddings to prevent knowledge and information loss during the training process. We additionally decode the final structural embeddings and the weighted sum of the two:
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>C</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The ConvE decoder is trained with binary cross-entropy loss:
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mi>L</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>t</mml:mi></mml:math></inline-formula> is the label vector, the elements of <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>t</mml:mi></mml:math></inline-formula> are 1 for relations existing and 0 otherwise.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Additional Missing Modality Completion</title>
<p>To address the issue of modality missing, we propose the <bold><italic>additional missing modality completion</italic></bold> mechanism, which primarily consists of two parts: <bold><italic>missing modality generation</italic></bold> and <bold><italic>missing modality imagination</italic></bold>. Firstly, we consider leveraging the powerful natural language understanding capabilities of large language models to extract semantic knowledge from textual information of entities and generate simulated multi-modal joint embeddings to replace the missing modalities. Additionally, inspired by VAE [<xref ref-type="bibr" rid="ref-32">32</xref>] and UMAEA [<xref ref-type="bibr" rid="ref-33">33</xref>], we train a variational multi-modal auto-encoder to generate potential probability distributions of the missing modalities, thereby assisting in the training process. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the architecture of the additional missing modality completion.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The overall architecture of the missing modality completion</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-4.tif"/>
</fig>
<sec id="s4_6_1">
<label>4.6.1</label>
<title>Missing Modality Generation</title>
<p>Given the powerful natural language understanding capabilities of large language models, we leverage the latent multi-modal knowledge embedded in multi-modal LLMs to address the missing modality issue in MMKGC. To be specific, we employ CLIP [<xref ref-type="bibr" rid="ref-34">34</xref>] as the missing modality generator. CLIP is a multi-modal large model that utilizes 400 million image-text pairs for contrastive training. Through large-scale training, multi-modal semantic information is implicitly stored in its model parameters.</p>
<p>For a modality missing entity <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, we extract its entity names and text descriptions, concatenate them, and then truncate the result for use as input <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to ensure it does not exceed the maximum input length of CLIP. CLIP first converts it to corresponding tokens:
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>d</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo stretchy="false">[</mml:mo><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mo stretchy="false">[</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mo stretchy="false">[</mml:mo><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> are special tokens representing the beginning and end of the text, respectively. The embedding of <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mo stretchy="false">[</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>S</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is usually used to represent the entire text semantic information. We use the text encoder in CLIP to directly obtain the input text features, and further obtain multi-modal joint embedding as virtual visual embedding <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> through normalization:
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>L</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>C</mml:mi><mml:mi>L</mml:mi><mml:mi>I</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>d</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The obtained virtual visual embeddings <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> will replace the original visual embeddings for entities that lack visual modalities.</p>
</sec>
<sec id="s4_6_2">
<label>4.6.2</label>
<title>Missing Modality Imagination</title>
<p>Inspired by VAE and UMAEA, based on a variational multi-modal auto-encoder framework, we train a set of Multi-Layer Perceptron (MLP) as encoder-decoder, input the original structural embedding <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, outputs the generated structural embedding <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. We adopt the hidden layer between them as another virtual visual feature <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2295;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>z</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi mathvariant="script">N</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the mean and variance of the simulated Gaussian distribution <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>z</mml:mi></mml:math></inline-formula>, respectively. And <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> is the generated structural embedding output by decoder.</p>
<p>In order to improve the quality of generated virtual visual features, we set three different loss functions to train the MLPs. Firstly, the KL divergence <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is minimized to train the variational autoencoder with a potential space that is close to a Gaussian distribution. Secondly, the differences between the generated virtual visual and structural features and the real features are minimized:
<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>r</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-27"><label>(27)</label><mml:math id="mml-eqn-27" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experiment</title>
<p>In this chapter, we will report the experiment details including datasets, evaluation index, parameter settings, baselines and the results. We conduct experiments to answer the following three questions about MMCSD:
<list list-type="order">
<list-item>
<p>Question 1 (Q1): Has MMCSD shown improvement compared to the previous baselines?</p></list-item>
<list-item>
<p>Question 2 (Q2): Compared to other models that focus on the issue of missing modalities, does our proposed missing modality completion mechanism exhibit superiority?</p></list-item>
<list-item>
<p>Question 3 (Q3): Is the design of each part of MMCSD valid?</p></list-item>
</list></p>
<sec id="s5_1">
<label>5.1</label>
<title>Datasets</title>
<p>We employed two multi-modal knowledge completion datasets for training and evaluation: the public benchmark FB15K-237, which is widely used in MMKGC, and DB13K, which we constructed based on knowledge from MMKB [<xref ref-type="bibr" rid="ref-3">3</xref>]. The specific statistics of the datasets are shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>The statistics of datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Entities</th>
<th>Relations</th>
<th>Train</th>
<th>Valid</th>
<th>Test</th>
<th>Multi-modal prop.</th>
</tr>
</thead>
<tbody>
<tr>
<td>FB15K-237</td>
<td>14541</td>
<td>237</td>
<td>272115</td>
<td>17535</td>
<td>20466</td>
<td>91.05%</td>
</tr>
<tr>
<td>DB13K</td>
<td>12842</td>
<td>279</td>
<td>59434</td>
<td>19800</td>
<td>19796</td>
<td>99.96%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>MMKB is a comprehensive multi-modal knowledge graph that encapsulates triples and visual data sourced from esteemed knowledge graphs like FB15K, DBpedia15K, and Yago15K. Inspired by the data format and structure of the mainstream MMKGC dataset FB15K-237, we have reconstructed the triples and visual modalities of DBpedia15K. We reassigned unique identifiers to the entities and relations in DBpedia15K and structured the triples accordingly. Additionally, we extracted pre-processed visual features associated with these entities and organized them into a visual feature matrix with the ordering determined by the newly assigned entity identifiers. This approach ensures consistency between visual features and triples in the newly created dataset, thereby enhancing its usability and versatility.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Evaluation Index</title>
<p>Following mainstream MMKG research, We use <italic>Hit@n</italic>, and Mean Reciprocal Rank (<italic>MRR</italic>) to objectively evaluate the effectiveness of the model. The larger <italic>Hit@n</italic> and <italic>MRR</italic> indicate the better performance of the model:
<disp-formula id="eqn-28"><label>(28)</label><mml:math id="mml-eqn-28" display="block"><mml:mi>H</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-29"><label>(29)</label><mml:math id="mml-eqn-29" display="block"><mml:mi>M</mml:mi><mml:mi>R</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p><italic>Hit@n</italic> represents the probability that the top <italic>n</italic> items of the candidate triple prediction possibility rank have correct results, and <italic>MRR</italic> represents the average of the reciprocal of correct ranking in the candidates.</p>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Implementation Details</title>
<p>MMCSD is trained on single NVIDIA RTX 3090 GPU. The proposed approach was implemented using Python 3.8.16, PyTorch 1.12.0, CUDA 12.1. We follow the step-by-step training method of KBAT, first training SpGAT to obtain embeddings of entities, relations, and visual information, then training the ConvE decoder for specific knowledge completion tasks.</p>
<p>The number of SpGAT training epochs is 3000, while for ConvE it is 300. We use Adam to optimize all the parameters with initial learning rate set at 0.001. Both the entity and relation embeddings of the final SpGAT layer are set to 200. The activation functions used for training ConvE are <italic>Sigmoid</italic> and <italic>ReLU</italic>.</p>
<p>For visual embedding, we use pre-processed visual features by trained VGG16 as original visual features [<xref ref-type="bibr" rid="ref-3">3</xref>]. We use the CLIP with ViT-B/32 as the multi-modal large model in MMCSD.</p>
<p>We follow the training approach of the UMAEA for missing modality imagination. In the first 1500 epochs of training SpGAT, only <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> was used for training, without involving missing modality imagination mechanisms. In the latter 1500 epochs, the modality imagination mechanism is added and SpGAT be trained using <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>. Because the missing modality imagination mechanism is used after multi-modal large model generation, we assume that all entities already contain their multi-modal information (including virtual multi-modal information), and during the decoding process, we weight the original embedding onto the final embedding. Therefore, we do not freeze the main model during training missing modality imagination mechanisms.</p>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Baselines</title>
<p>We compared our MMCSD against several knowledge graph completion models, encompassing both uni-modal and multi-modal approaches to demonstrate the superiority of our MMCSD. Firstly, we choose the conventional text-based uni-modal models for comparison to demonstrate the improvement brought by the visual information. For uni-modal models, we included TransE [<xref ref-type="bibr" rid="ref-11">11</xref>] and RotatE [<xref ref-type="bibr" rid="ref-13">13</xref>], which consider the relation in triples as translation vector from head entities to tail entities in different vector spaces. Additionally, we considered DistMult [<xref ref-type="bibr" rid="ref-35">35</xref>] and ComplEx [<xref ref-type="bibr" rid="ref-36">36</xref>], which attach great importance to mining the potential semantics of entities and relations. We also considered neural network-based models such as KBAT [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>Secondly, we compared MMCSD with multi-modal models, including TransAE [<xref ref-type="bibr" rid="ref-3">3</xref>] and IKRL [<xref ref-type="bibr" rid="ref-1">1</xref>], which respectively extend TransE based on visual representations and multi-modal AutoEncoder. VBKGC [<xref ref-type="bibr" rid="ref-23">23</xref>] employs VisualBERT as a multi-modal encoder to capture the deeply fused multi-modal features of entities. MKBE [<xref ref-type="bibr" rid="ref-22">22</xref>] uses uni-modal embeddings as the context for attribute-specific decoders to generate the missing values of triples. MKGformer [<xref ref-type="bibr" rid="ref-37">37</xref>] leverages a hybrid transformer architecture with unified input-output and performs multi-modal fusion in multi-level. MoSE [<xref ref-type="bibr" rid="ref-5">5</xref>] learns modality-split embeddings for each modality to alleviate the modality interference. IMF [<xref ref-type="bibr" rid="ref-38">38</xref>] employs a two-stage multi-modal fusion framework to preserve modality-specific knowledge as well as to take advantage of the complementarity between different modalities. LAFA [<xref ref-type="bibr" rid="ref-39">39</xref>] designs a modality interaction attention mechanism that dynamically measures the contribution of images to entity embedding based on link information, thereby mitigating the impact of extraneous information in the visual modality on the complementation effect. CMR [<xref ref-type="bibr" rid="ref-40">40</xref>] proposes a unified cross-modal contrast learning to simultaneously capture multi-modal correlations of query-entity pairs in a unified representation space to improve the similarity of representations of useful semantic neighbors and then support the semantic neighbor retrieval. HKA [<xref ref-type="bibr" rid="ref-41">41</xref>] proposes macro- and micro-knowledge alignment module to capture the global semantic relevance between modalities and more effectively reveal the local consistency information through multi-modal supervisory effects.</p>
<p>In addition, we also compared MMCSD with MACO [<xref ref-type="bibr" rid="ref-7">7</xref>], which also considers the problem of modality missing and proposes to generate auxiliary embeddings by generative adversarial networks, to analyze the effectiveness of the additional missing modality completion method.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Main Results and Analysis (Q1 &#x0026; Q2)</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> shows the main results of the MMKGC experiments on two datasets (&#x002A;: There is a label leakage error in KBAT, so the corrected result is poor compared with the result in their paper [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]), some baselines were not tested on the DB13K due to their code not being open-source or requiring additional external knowledge.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparative results of MMCSD against other baseline methods on two datasets</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Models</th>
<th colspan="3">FB15K-237</th>
<th colspan="3">DB13K</th>
</tr>
<tr>
<th>Hits@1</th>
<th>Hits@3</th>
<th>MRR</th>
<th>Hits@1</th>
<th>Hits@3</th>
<th>MRR</th>
</tr>
</thead>
<tbody>
<tr>
<td>TransE</td>
<td>0.198</td>
<td>0.376</td>
<td>0.279</td>
<td>0.155</td>
<td>0.390</td>
<td>0.292</td>
</tr>
<tr>
<td>KBAT&#x002A;</td>
<td>0.183</td>
<td>0.317</td>
<td>0.287</td>
<td>0.177</td>
<td>0.283</td>
<td>0.249</td>
</tr>
<tr>
<td>DistMult</td>
<td>0.199</td>
<td>0.301</td>
<td><bold>&#x2013;</bold></td>
<td>0.185</td>
<td>0.299</td>
<td>0.256</td>
</tr>
<tr>
<td>ComplEx</td>
<td>0.194</td>
<td>0.297</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>RotatE</td>
<td>0.241</td>
<td>0.375</td>
<td>0.338</td>
<td>0.228</td>
<td>0.358</td>
<td>0.317</td>
</tr>
<tr>
<td>TransAE</td>
<td>0.199</td>
<td>0.317</td>
<td><bold>&#x2013;</bold></td>
<td>0.161</td>
<td>0.376</td>
<td>0.289</td>
</tr>
<tr>
<td>IKRL</td>
<td>0.194</td>
<td>0.284</td>
<td><bold>&#x2013;</bold></td>
<td>0.159</td>
<td>0.370</td>
<td>0.282</td>
</tr>
<tr>
<td>VBKGC</td>
<td>0.213</td>
<td>0.332</td>
<td>0.301</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>MKBE</td>
<td>0.258</td>
<td><bold>&#x2013;</bold></td>
<td>0.347</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>MKGformer</td>
<td>0.256</td>
<td>0.367</td>
<td><bold>&#x2013;</bold></td>
<td>0.241</td>
<td>0.369</td>
<td>0.323</td>
</tr>
<tr>
<td>MoSE</td>
<td>0.281</td>
<td>0.411</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>IMF</td>
<td>0.287</td>
<td><bold>&#x2013;</bold></td>
<td>0.389</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>LAFA</td>
<td>0.269</td>
<td>0.398</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>CMR</td>
<td>0.263</td>
<td>0.395</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>HKA</td>
<td>0.291</td>
<td>0.424</td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
<td><bold>&#x2013;</bold></td>
</tr>
<tr>
<td>MMCSD (ours)</td>
<td><bold>0.294</bold><sup><bold>1</bold></sup></td>
<td><bold>0.425</bold></td>
<td><bold>0.395</bold></td>
<td><bold>0.291</bold></td>
<td><bold>0.405</bold></td>
<td><bold>0.384</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: <sup><bold>1</bold></sup>Bold font in the table indicates the best results. &#x002A;There is a label leakage error in KBAT, so the corrected result is poor compared with the result in their paper [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>].</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The comparison of results between MMCSD and uni-modal models in <xref ref-type="table" rid="table-2">Table 2</xref> demonstrates the importance of visual modality in MMKGC. Compared to uni-modal models, the majority of models that incorporate visual modalities show some degree of performance improvement. However, it is noticeable that the use of multi-modal information does not necessarily enhance the effectiveness of knowledge completion. For instance, some indices for IKRL and TransAE are even inferior to those of TransE. We believe that the poor performance of some models that add visual modalities may be attributed to their use of rough visual information and primitive modal fusion methods.</p>

<p>Compared with multi-modal models, MMCSD achieved significant improvement in all indices. Specifically, compared to the best baseline IMF, <italic>Hits@1</italic> has increased by 0.7%, and <italic>MRR</italic> has improved by 0.6%. This demonstrates the positive effect of deep semantic exploration in various modalities and missing modality completion on MMKGC.</p>
<p>On our self-constructed DB13K, we similarly obtained comparable results to those from the FB15K-237, proving the generalizability of DB13K. Although the training set of DB13K is relatively small, it is difficult to obtain more supervised information compared to FB15K-237. However, due to the superiority of its data format and structure and the high richness of multi-modal information, the experiments conducted on it still achieved the same results as we expected.</p>
<p>To verify the effectiveness of missing modality completion on MMKGC, we set the proportion of multi-modal information in FB15K-237 to 20%, 40%, 60%, and 80% for experiments, while comparing MMCSD which only used additional missing modality completion (MMC-only) with MACO which employed IKRL as score function (IKRL-based). The experimental results for MACO, as reported in <xref ref-type="fig" rid="fig-3">Fig. 3</xref> of their paper [<xref ref-type="bibr" rid="ref-7">7</xref>], due to the lack of specific numerical values, only display the approximate maximum value. The contrast results are shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison results of MMCSD which only used additional missing modality completion (MMC-only) with MACO which employed IKRL as score function (IKRL-based) in FB15K-237 with different multi-modal ratios</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-5.tif"/>
</fig>
<p>It can be seen that compared to MACO, significant improvements have been made. Due to the different encoder-decoder and multi-modal fusion methods used, it cannot definitively be concluded that the missing modality generation in MMCSD is superior to the mechanism used in MACO. However, using pre-trained multi-modal large models and training VAE can save computational resources more, which is one of the advantages of our proposed method. It is worth noting that the completeness of the original visual features does not necessarily correlate with the effectiveness of MMKGC. The specific promotion mechanism of visual modality for MMKGC is still not fully understood. Therefore, the methods and timing for integrating multi-modal information, as well as its specific impact on knowledge completion, should be key areas of focus in our future research.</p>
<p>To evaluate the robustness of MMCSD, we conducted experiments on the FB15K-237 where entity description images were subjected to adversarial attacks. Considering that the original visual features were extracted by VGG16, we performed the Projected Gradient Descent (PGD) attack on VGG16 trained on the CIFAR10 targeting the image classification task and subsequently used the attacked model to generate adversarial samples. The PGD is implemented based on the open source Python library Adversarial Robustness Toolbox (ART).</p>
<p>The experimental results are presented in <xref ref-type="table" rid="table-3">Table 3</xref>. It can be seen that MMCSD has better robustness compared to the model that only employs multi-modal fusion. By analyzing some of the cases, we find that although the original visual features of the described image are greatly changed after the attack, the accuracy of detail extraction by mature models such as object recognition and optical character recognition is not significantly reduced, so that the quality of Multi-level visual Semantic Extraction and Triple Semantic Embedding is still guaranteed.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparative results of MMCSD on FB15K-237 after adversarial attacks</title>
</caption>
<table>
<colgroup>
<col width="30mm"/>
<col width="20mm"/>
<col width="20mm"/>
<col width="20mm"/>
<col width="20mm"/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Datasets</th>
<th colspan="2">MMCSD (only modal fusion)</th>
<th colspan="2">MMCSD</th>
</tr>
<tr>
<th>Hits@1</th>
<th>MRR</th>
<th>Hits@1</th>
<th>MRR</th>
</tr>
</thead>
<tbody>
<tr>
<td>FB15K-237</td>
<td>0.215</td>
<td>0.307</td>
<td>0.294</td>
<td>0.395</td>
</tr>
<tr>
<td>FB15K-237 (B.A.)</td>
<td>0.198</td>
<td>0.288</td>
<td>0.283</td>
<td>0.386</td>
</tr>
<tr>
<td><italic>Decline</italic></td>
<td><italic>0.017</italic></td>
<td><italic>0.019</italic></td>
<td><italic>0.011</italic></td>
<td><italic>0.009</italic></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Ablation Study (Q3)</title>
<p>The ablation study is further conducted to prove the effects of different improvement methods. We designed experiments to compare the models without Super-Resolution (w/o SR), without Missing Modality Completion (w/o MMC), without Multi-level visual Semantic Extraction (w/o MSE), and without Triple Semantic Embedding (w/o TSE) respectively.</p>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, the improvement effect of adopting MMC on the model is relatively weak, possibly because at least 91% of the entities in FB15K-237 contain descriptive image information, and other mechanisms also have a certain complementary effect on semantics. Meanwhile, the comparison of the performance with models that only applied multi-modal fusion also proves the effectiveness of enhancement methods proposed in this paper for MMKGC.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>The results of the ablation experiment</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Models</th>
<th colspan="3">FB15K-237</th>
</tr>
<tr>
<th>Hits@1</th>
<th>Hits@3</th>
<th>MRR</th>
</tr>
</thead>
<tbody>
<tr>
<td>MMCSD (only modal fusion)</td>
<td>0.215</td>
<td>0.329</td>
<td>0.307</td>
</tr>
<tr>
<td>MMCSD (w/o MSE)</td>
<td>0.272</td>
<td>0.405</td>
<td>0.359</td>
</tr>
<tr>
<td>MMCSD (w/o SR)</td>
<td>0.279</td>
<td>0.405</td>
<td>0.364</td>
</tr>
<tr>
<td>MMCSD (w/o TSE)</td>
<td>0.278</td>
<td>0.410</td>
<td>0.362</td>
</tr>
<tr>
<td>MMCSD (w/o MMC)</td>
<td>0.287</td>
<td>0.416</td>
<td>0.388</td>
</tr>
<tr>
<td>MMCSD</td>
<td><bold>0.294</bold><sup><bold>1</bold></sup></td>
<td><bold>0.425</bold></td>
<td><bold>0.395</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-4fn1" fn-type="other">
<p>Note: <sup><bold>1</bold></sup>Bold font in the table indicates the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Except for Missing Modality Completion, the other three mechanisms have effectively improved various indicators of MMKGC. Among them, Multi-level Visual Semantic Extraction, which aims at exploring semantics in visual modalities, has the best performance, consistent with our expectations. Visual modality, as the most unique and semantically rich information in MMKGs, has a significant promoting effect on the representation of multi-modal entities and triples. Therefore, further analysis and processing of visual modalities should be the key to improving the effectiveness of MMKGC in our future work.</p>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows the decrease in indices after removing various optimization methods. The greater the decrease, the more important the optimization method is for improving the model performance. MSE, SR, and TSE add more semantic information of entities from different levels, which obviously improves the effectiveness of knowledge completion. MSE and SR respectively optimized the detail analysis mechanism and global semantic features of entity images, which have better improvement effects compared to the TSE method using triple text semantics. The MMC method, as a supplement and enhancement for modality missing entities, has a certain effect, but due to the limited number of modality missing entities in FB15K-237, the optimization is not as high as the other three methods.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The impact of removing various optimization methods on MMCSD</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_60395-fig-6.tif"/>
</fig>
</sec>
<sec id="s5_7">
<label>5.7</label>
<title>Complexity Analysis</title>
<p>The whole model implementation can be divided into two phases: feature pre-extraction and model training. The feature pre-extraction phase involves the acquisition of the original visual features (usually the original visual features extracted by VGG16 are already available in the dataset), the super-resolution of the original image, the multi-level visual semantic extraction, the generation of triple semantic embeddings, and the implementation of the missing modality generation mechanism. The model training phase includes the training of the graph attention network encoder, the multi-modal fusion layer, the decoder, and the missing modality imagination mechanism.</p>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows the performance and efficiency of the model training phase of MMCSD. Compared to the vanilla uni-modal knowledge completion model, MMCSD only adds the multi-modal fusion layer and the variational multi-modal auto-encoder, and therefore does not require significantly more computation and time in the model training phase.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance evaluation metrics for MMCSD</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Datasets</th>
<th>Avg. training time per epoch</th>
<th>FLOPs</th>
<th>Params.</th>
</tr>
</thead>
<tbody>
<tr>
<td>FB15k-237</td>
<td>3.458 s</td>
<td>16.1964 G</td>
<td rowspan="2">1.1108 M</td>
</tr>
<tr>
<td>DB13K</td>
<td>1.326 s</td>
<td>14.3258 G</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As we use a number of well-established techniques to extract semantic information from the original multi-modal data, some additional time is required beyond the model training phase. <xref ref-type="table" rid="table-6">Table 6</xref> shows the average time to process a single piece of data for the established methods covered in this paper when using a single 3090 GPU, and the application of large language model to summarize the semantics of entities extracted by other methods and generate descriptive text takes more time and computational resources. As the dataset expands, the time required for the feature pre-extraction phase will increase accordingly. The total time depends only on the number of entities in the dataset and the number of descriptive images, but not on the number of triples.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Average processing time per data item (text or image) for the each mature model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Models</th>
<th>Average time (s)</th>
<th>GPU memory usage (MiB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>ESDR</td>
<td>2.075</td>
<td>3690</td>
</tr>
<tr>
<td>BLIP2</td>
<td>2.406</td>
<td>15360</td>
</tr>
<tr>
<td>PPOCR</td>
<td>3.686</td>
<td>1460</td>
</tr>
<tr>
<td>GRiT</td>
<td>0.21</td>
<td>6340</td>
</tr>
<tr>
<td>FastSAM</td>
<td>0.052</td>
<td>3080</td>
</tr>
<tr>
<td>Qwen1.5</td>
<td>15.478</td>
<td>16450</td>
</tr>
<tr>
<td>CLIP</td>
<td>1.972</td>
<td>1808</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>Aiming at the problem of low modality quality and modality missing, we propose MMCSD, a multi-modal knowledge graph completion model based on super-resolution and detailed description generation. The core idea is to use super-resolution and multi-level image semantic extraction to further enrich semantic information of multi-modal embeddings. Moreover, we additionally use the missing modality generation and imagination mechanisms to complete the visual feature of modality missing entities. The experiment shows that our proposed methods have a significant positive effect on MMKGC.</p>
<p>In future work, we will attempt to use more novel graph neural networks, transformer architecture, and even large language models to extract deep information in graph-structured data to build better-quality structural embeddings. Secondly, it can be seen from the experiment that the promotion effect of multi-modal information and its fusion timing and method on MMKGC is not clear enough. We will strive to explore its principles and further optimize the multi-modal knowledge completion effect.</p>
</sec>
</body>
<back>
<ack>
<p>The authors are grateful to all the editors and reviewers for their detailed review and insightful advice.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Research Project, grant number BHQ090003000X03.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Huansha Wang, Ruiyang Huang; data collection: Huansha Wang; analysis and interpretation of results: Huansha Wang, Qinrang Liu; draft manuscript preparation: Huansha Wang; visualization: Shaomei Li, Jianpeng Zhang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are available from the corresponding author, Huansha Wang, upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>MMKG</term>
<def>
<p>Multi-modal knowledge graph</p>
</def>
</def-item>
<def-item>
<term>MMKGC</term>
<def>
<p>Multi-modal knowledge graph completion</p>
</def>
</def-item>
<def-item>
<term>MMCSD</term>
<def>
<p>Multi-modal knowledge graph Completion model based on Super-resolution and Detailed Description Generation</p>
</def>
</def-item>
<def-item>
<term>LLM</term>
<def>
<p>Large language model</p>
</def>
</def-item>
<def-item>
<term>GCN</term>
<def>
<p>Graph convolutional network</p>
</def>
</def-item>
<def-item>
<term>GAT</term>
<def>
<p>Graph attention network</p>
</def>
</def-item>
<def-item>
<term>SpGAT</term>
<def>
<p>Sparse graph attention network</p>
</def>
</def-item>
<def-item>
<term>PVM</term>
<def>
<p>Pre-trained visual model</p>
</def>
</def-item>
<def-item>
<term>MLP</term>
<def>
<p>Multi-Layer Perceptron</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Luan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Image-embodied knowledge representation learning</article-title>. In: <conf-name>Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence</conf-name>; <year>2017 Aug 19&#x2013;26</year>; <publisher-loc>Melbourne, VIC, Australia</publisher-loc>: <publisher-name>International Joint Conferences on Artificial Intelligence Organization</publisher-name>. Vol. <volume>2017</volume>, p. <fpage>3140</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2017/438</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mousselly Sergieh</surname> <given-names>H</given-names></string-name>, <string-name><surname>Botschen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gurevych</surname> <given-names>I</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name></person-group>. <article-title>A multimodal translation-based approach for knowledge Graph Representation learning</article-title>. In: <conf-name>Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics</conf-name>; <year>2018</year>; <publisher-loc>New Orleans, Louisiana, Stroudsburg, PA, USAACL</publisher-loc>. p. <fpage>225</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/s18-2027</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Multimodal data enhanced representation learning for knowledge graphs</article-title>. In: <conf-name>2019 International Joint Conference on Neural Networks (IJCNN)</conf-name>; <year>2019 Jul 14&#x2013;19</year>; <publisher-loc>Budapest, Hungary</publisher-loc>: <publisher-name>IEEE</publisher-name>. Vol. <volume>2019</volume>, p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ijcnn.2019.8852079</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Modality-aware negative sampling for multi-modal knowledge graph embedding</article-title>. In: <conf-name>2023 International Joint Conference on Neural Networks (IJCNN)</conf-name>; <year>2023 Jun 18&#x2013;23</year>; <publisher-loc>Gold Coast, QSL, Australia</publisher-loc>: <publisher-name>IEEE</publisher-name>. Vol. <volume>2023</volume>, p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/IJCNN54540.2023.10191314</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MoSE: modality split and ensemble for multimodal knowledge graph completion</article-title>. In: <conf-name>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</conf-name>; <year>2022</year>; <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>. p. <fpage>10527</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2022.emnlp-main</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Making large language models perform better in knowledge graph completion</article-title>. In: <conf-name>Proceedings of the 32nd ACM International Conference on Multimedia</conf-name>; <year>2024</year>; <publisher-loc>Melbourne VIC Australia</publisher-loc>: <publisher-name>ACM</publisher-name>. p. <fpage>233</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3664647.3681327</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name></person-group>. <chapter-title>MACO: a modality adversarial and contrastive framework for modality-missing multi-modal knowledge graph completion</chapter-title>. In: <source>Natural language processing and Chinese computing</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2023</year>. p. <fpage>123</fpage>&#x2013;<lpage>34</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-031-44693-1_10</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lim</surname> <given-names>B</given-names></string-name>, <string-name><surname>Son</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>KM</given-names></string-name></person-group>. <article-title>Enhanced deep residual networks for single image super-resolution</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>; <year>2017 Jul 21&#x2013;26</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. Vol. <volume>2017</volume>, p. <fpage>1132</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW.2017.151</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Pantel</surname> <given-names>P</given-names></string-name>, <string-name><surname>Poon</surname> <given-names>H</given-names></string-name>, <string-name><surname>Choudhury</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gamon</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Representing text for joint embedding of text and knowledge bases</article-title>. In: <conf-name>Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</conf-name>; <year>2015</year>; <publisher-loc>Lisbon, Portugal, Stroudsburg, PA, USAACL</publisher-loc>; p. <fpage>1499</fpage>&#x2013;<lpage>509</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/d15-1174</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Garcia-Duran</surname> <given-names>A</given-names></string-name>, <string-name><surname>Niepert</surname> <given-names>M</given-names></string-name>, <string-name><surname>Onoro-Rubio</surname> <given-names>D</given-names></string-name>, <string-name><surname>Rosenblum</surname> <given-names>DS</given-names></string-name></person-group>. <chapter-title>MMKG: multi-modal knowledge graphs</chapter-title>. In: <source>The semantic web</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2019</year>. p. <fpage>459</fpage>&#x2013;<lpage>74</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-030-21348-0_30</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bordes</surname> <given-names>A</given-names></string-name>, <string-name><surname>Usunier</surname> <given-names>N</given-names></string-name>, <string-name><surname>Duran</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Weston</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yakhnenko</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Translating embeddings for modeling multi relational data</article-title>. In: <conf-name>NIPS&#x2019;13: Proceedings of the 27th International Conference on Neural Information Processing Systems (NIPS)</conf-name>; <year>2013 Dec 5&#x2013;8</year>; <publisher-loc>Lake Tahoe, Nevada, USA</publisher-loc>. p. <fpage>2787</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Knowledge graph embedding by translating on hyperplanes</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2014</year>:<fpage>1112</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v28i1.8870</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>RotatE: knowledge graph embedding by relational rotation in complex space</article-title>. <comment>arXiv:1902.10197. 2019</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dettmers</surname> <given-names>T</given-names></string-name>, <string-name><surname>Minervini</surname> <given-names>P</given-names></string-name>, <string-name><surname>Stenetorp</surname> <given-names>P</given-names></string-name>, <string-name><surname>Riedel</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Convolutional 2D knowledge graph embeddings</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2018</year>;<fpage>1811</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v32i1.11573</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>DQ</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>TD</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>DQ</given-names></string-name>, <string-name><surname>Phung</surname> <given-names>D</given-names></string-name></person-group>. <article-title>A novel embedding model for knowledge base completion based on Convolutional neural network</article-title>. In: <conf-name>Proceedings of the 2018 Conference of the North American Chapter Of the Association for Computational Linguistics: Human Language Technologies</conf-name>; <year>2018</year>; <publisher-loc>New Orleans, LA, USA</publisher-loc>; p. <fpage>327</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/n18-2053</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Schlichtkrull</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kipf</surname> <given-names>TN</given-names></string-name>, <string-name><surname>Bloem</surname> <given-names>P</given-names></string-name>, <string-name><surname>van den Berg</surname> <given-names>R</given-names></string-name>, <string-name><surname>Titov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <chapter-title>Modeling relational data with graph convolutional networks</chapter-title>. In: <source>The semantic web</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2018</year>. p. <fpage>593</fpage>&#x2013;<lpage>607</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-319-93417-4_38</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nathani</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chauhan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kaul</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Learning attention-based embeddings for relation prediction in knowledge graphs</article-title>. In: <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</conf-name>; <year>2019</year>; <publisher-loc>Florence, Italy, Stroudsburg, PA, USAACL</publisher-loc>; p. <fpage>4710</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/p19-1466</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Attention is all you need</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>; <year>2017 Dec 4&#x2013;9</year>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Ferrada</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bustos</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hogan</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>IMGpedia: a linked dataset with content-based analysis of wikimedia images</chapter-title>. In: <source>The semantic web-ISWC 2017</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2017</year>. p. <fpage>84</fpage>&#x2013;<lpage>93</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-319-68204-4_8</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Alberts</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Deshpande</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Vania</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>VisualSem: a high-quality knowledge graph for vision and language</article-title>. In: <conf-name>Proceedings of the 1st Workshop on Multilingual Representation Learning</conf-name>; <year>2021</year>; <publisher-loc>Punta Cana, Dominican Republic. Stroudsburg, PA, USAACL</publisher-loc>. p. <fpage>138</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2021.mrl-1.13</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>L</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GAKG: a multimodal geoscience academic knowledge graph</article-title>. In: <conf-name>Proceedings of the 30th ACM International Conference on Information &#x0026; Knowledge Management</conf-name>; <year>2021 Nov 1&#x2013;5</year>; <publisher-loc>Queensland, Australia</publisher-loc>. p. <fpage>4445</fpage>&#x2013;<lpage>54</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3459637.3482003</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pezeshkpour</surname> <given-names>P</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Embedding multimodal relational data for knowledge base completion</article-title>. In: <conf-name>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</conf-name>; <year>2018</year>; <publisher-loc>Brussels, Belgium, Stroudsburg, PA, USAACL</publisher-loc>. p. <fpage>3208</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/d18-1359</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Knowledge graph completion with pre-trained multimodal transformer and twins negative sampling</article-title>. <comment>arXiv:2209.07084. 2022</comment>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kipf</surname> <given-names>TN</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Semi supervised classification with graph convolutional networks</article-title>. <comment>arXiv:1609.02907. 2017</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Velickovic</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cucurull</surname> <given-names>G</given-names></string-name>, <string-name><surname>Casanova</surname> <given-names>A</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li&#x000F2;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Graph attention networks</article-title>. <comment>arXiv:1710.10903. 2018</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Monkey: image resolution and text label are important things for large multi-modal models</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <year>2024 Jun 16&#x2013;22</year>; <publisher-loc>Seattle, WA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. Vol. <volume>2024</volume>, p. <fpage>26753</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52733.2024.02527</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Savarese</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>SCH</given-names></string-name></person-group>. <article-title>BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2023 Jul 23&#x2013;29</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>19730</fpage>&#x2013;<lpage>42</lpage>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>PP-OCR: a practical ultra lightweight OCR system</article-title>. <comment>arXiv:2009.09941. 2020</comment>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GRiT: a generative region-to-text transformer for object understanding</article-title>. In: <conf-name>Computer Vision-ECCV 2024-18th European Conference</conf-name>; <year>2024 Sep 29&#x2013;Oct 4</year>; <publisher-loc>Milan, Italy</publisher-loc>. p. <fpage>207</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-72989-8_12</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>W</given-names></string-name>, <string-name><surname>An</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fast segment anything</article-title>. <comment>arXiv:2306.12156. 2023</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Dang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Qwen technical report</article-title>. <comment>arXiv.2309.16609. 2024</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational bayes</article-title>. <comment>arXiv:1312.6114. 2022</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Rethinking uncertainly missing and ambiguous visual modality in multi-modal entity alignment</article-title>. In: <conf-name>22nd International Semantic Web Conference</conf-name>; <year>2023 Nov 6&#x2013;10</year>; <publisher-loc>Athens, Greece</publisher-loc>; p. <fpage>121</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-47240-4</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Learning transferable visual models from natural language supervision</article-title>. In: <conf-name>Proceedings of the 38th International Conference on Machine Learning (PMLR 139)</conf-name>; <year>2021 Jul 18&#x2013;24</year>; p. <fpage>8748</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yih</surname> <given-names>W</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Embedding entities and relations for learning and inference in knowledge bases</article-title>. In: <conf-name>3rd International Conference on Learning Representations, ICLR 2015</conf-name>; <year>2015 May 7&#x2013;9</year>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Trouillon</surname> <given-names>T</given-names></string-name>, <string-name><surname>Welbl</surname> <given-names>J</given-names></string-name>, <string-name><surname>Riedel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gaussier</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bouchard</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Complex embeddings for simple link prediction</article-title>. In: <conf-name>Proceedings of the 33rd International Conference on Machine Learning</conf-name>; <year>2016 Jun 19&#x2013;24</year>; <publisher-loc>New York City, NY, USA</publisher-loc>. p. <fpage>2071</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Hybrid transformer with multi-level fusion for multimodal knowledge graph completion</article-title>. In: <conf-name>Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>; <year>2022</year>; <publisher-loc>Madrid Spain</publisher-loc>: <publisher-name>ACM</publisher-name>. p. <fpage>904</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3477495.3531992</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>C</given-names></string-name></person-group>. <article-title>IMF: interactive multimodal fusion model for link prediction</article-title>. In: <conf-name>Proceedings of the ACM Web Conference 2023</conf-name>; <year>2023</year>; <publisher-loc>Austin TX USA</publisher-loc>: <publisher-name>ACM</publisher-name>. p. <fpage>2572</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3543507.3583554</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>D</given-names></string-name></person-group>. <article-title>LAFA: multimodal knowledge graph completion with link aware fusion and aggregation</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2024</year>;<volume>38</volume>(<issue>8</issue>):<fpage>8957</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i8.28744</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>B</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>X</given-names></string-name>, <string-name><surname>Song</surname> <given-names>K</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Contrast then memorize: semantic neighbor retrieval-enhanced inductive multimodal knowledge graph completion</article-title>. In: <conf-name>Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval</conf-name>; <year>2024</year>; <publisher-loc>Washington, DC, USA</publisher-loc>: <publisher-name>ACM</publisher-name>. p. <fpage>102</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3626772.3657838</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>HKA: a hierarknowledge alignment framework for multimodal knowledge graph completion</article-title>. <source>ACM Trans Multimed Comput Commun Appl</source>. <year>2024</year>;<volume>20</volume>(<issue>8</issue>):<fpage>1</fpage>&#x2013;<lpage>19</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3664288</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>