<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">66647</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2025.066647</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Anime Generation through Diffusion and Language Models: A Comprehensive Survey of Techniques and Trends</article-title>
<alt-title alt-title-type="left-running-head">Anime Generation through Diffusion and Language Models: A Comprehensive Survey of Techniques and Trends</alt-title>
<alt-title alt-title-type="right-running-head">Anime Generation through Diffusion and Language Models: A Comprehensive Survey of Techniques and Trends</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Wu</surname><given-names>Yujie</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Deng</surname><given-names>Xing</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>xdeng@just.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Shao</surname><given-names>Haijian</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Cheng</surname><given-names>Ke</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Ming</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Jiang</surname><given-names>Yingtao</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Fei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Computer Science, Jiangsu University of Science and Technology</institution>, <addr-line>Zhenjiang, 212003</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Electrical and Computer Engineering University of Nevada</institution>, <addr-line>Las Vegas, NV 89154</addr-line>, <country>USA</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Xing Deng. Email: <email>xdeng@just.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>30</day><month>09</month><year>2025</year>
</pub-date>
<volume>144</volume>
<issue>3</issue>
<fpage>2709</fpage>
<lpage>2778</lpage>
<history>
<date date-type="received">
<day>14</day>
<month>4</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_66647.pdf"></self-uri>
<abstract>
<p>The application of generative artificial intelligence (AI) is bringing about notable changes in anime creation. This paper surveys recent advancements and applications of diffusion and language models in anime generation, focusing on their demonstrated potential to enhance production efficiency through automation and personalization. Despite these benefits, it is crucial to acknowledge the substantial initial computational investments required for training and deploying these models. We conduct an in-depth survey of cutting-edge generative AI technologies, encompassing models such as Stable Diffusion and GPT, and appraise pivotal large-scale datasets alongside quantifiable evaluation metrics. Review of the surveyed literature indicates the achievement of considerable maturity in the capacity of AI models to synthesize high-quality, aesthetically compelling anime visual images from textual prompts, alongside discernible progress in the generation of coherent narratives. However, achieving perfect long-form consistency, mitigating artifacts like flickering in video sequences, and enabling fine-grained artistic control remain critical ongoing challenges. Building upon these advancements, research efforts have increasingly pivoted towards the synthesis of higher-dimensional content, such as video and three-dimensional assets, with recent studies demonstrating significant progress in this burgeoning field. Nevertheless, formidable challenges endure amidst these advancements. Foremost among these are the substantial computational exigencies requisite for training and deploying these sophisticated models, particularly pronounced in the realm of high-dimensional generation such as video synthesis. Additional persistent hurdles include maintaining spatial-temporal consistency across complex scenes and mitigating ethical considerations surrounding bias and the preservation of human creative autonomy. This research underscores the transformative potential and inherent complexities of AI-driven synergy within the creative industries. We posit that future research should be dedicated to the synergistic fusion of diffusion and autoregressive models, the integration of multimodal inputs, and the balanced consideration of ethical implications, particularly regarding bias and the preservation of human creative autonomy, thereby establishing a robust foundation for the advancement of anime creation and the broader landscape of AI-driven content generation.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Diffusion models</kwd>
<kwd>language models</kwd>
<kwd>anime generation</kwd>
<kwd>image synthesis</kwd>
<kwd>video generation</kwd>
<kwd>stable diffusion</kwd>
<kwd>AIGC</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62202210</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>At the forefront of this transformation are diffusion models and language models, two classes of generative AI designed to bridge textual and visual domains. These powerful models are significantly impacting certain creative sectors like anime generation and are increasingly relevant across diverse scientific and engineering fields [<xref ref-type="bibr" rid="ref-1">1</xref>]. Diffusion models have demonstrated strong capabilities in synthesizing high-quality images and videos [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>]. Studies have shown that these models have improved performance improved performance compared to traditional generative adversarial networks (GANs) in image synthesis in terms of general quality, diversity, and specific metrics like FID [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Meanwhile, Language models such as BERT and GPT have significantly advanced the field of natural language processing [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. The integration of these technologies has enabled systems to interpret textual prompts and generate corresponding anime-style visuals with notable fidelity, a capability demonstrated by models like Stable Diffusion [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>This survey focuses on pivotal diffusion and language models that have demonstrated substantial impact and are directly relevant to anime content generation. Specifically, this comprehensive review addresses the pivotal question: How are diffusion and language models currently advancing and transforming the landscape of anime content generation, and what are the key challenges and future directions in effectively leveraging these technologies? Our selection criteria prioritize models with demonstrated efficacy in synthesizing anime-style visuals, coherent narratives, and related creative assets, alongside their technical innovation and prominence within the generative AI landscape.</p>
<p>This study investigates the application of these models across key anime production domains: narrative and graphic novel genesis, illustrative and keyframe synthesis, and episodic and interactive media expansion. For instance, language models can automate script generation and dialogue creation, while diffusion models synthesize anime-stylized imagery and sequential frames [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. The integration of these technologies addresses enduring challenges in creative efficiency, content personalization, and fiscal optimization. While necessitating significant computational resources and capital outlay for high-performance infrastructure, these technologies can streamline certain manual workflows, potentially leading to time savings and reduced labor-intensive operational costs in anime production. This shift has led to notable advancements by streamlining complex processes and accelerating content generation in specific areas [<xref ref-type="bibr" rid="ref-11">11</xref>]. The use of large-scale datasets like LAION-5B enhances the multilingual and stylistic capabilities of these models [<xref ref-type="bibr" rid="ref-12">12</xref>], broadening their applicability. Furthermore, a synergistic analysis employing automated (e.g., Fr&#x00E9;chet Inception Distance, CLIP Score) metrics alongside human-centered evaluation provides a comprehensive paradigm for appraising the efficacy of sophisticated diffusion and language models. This approach elucidates their technical viability and observed capacity to foster creative innovation within animation production [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>], thereby elucidating both their technical viability and their capacity to foster creative innovation within animation production.</p>
<p>Despite their promise, these technologies face significant hurdles. Technical challenges include maintaining consistency across multi-frame sequences and multi-character scenes. Ethical considerations, such as originality, copyright, and the potential diminishment of human creativity, also loom large [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>The overall structure of this study is depicted in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>: <xref ref-type="sec" rid="s2">Section 2</xref> offers a comprehensive background on diffusion models and language models, detailing their theoretical foundations and development history. <xref ref-type="sec" rid="s3">Section 3</xref> explores image generation methodologies, focusing on Stable Diffusion and its ecosystem for anime-style synthesis. <xref ref-type="sec" rid="s4">Section 4</xref> examines video generation, addressing advancements in temporal consistency and character animation. <xref ref-type="sec" rid="s5">Section 5</xref> delves into music composition, while <xref ref-type="sec" rid="s6">Section 6</xref> investigates game generation, extending the application of these models to interactive media. <xref ref-type="sec" rid="s7">Section 7</xref> covers alternative applications, such as narrative synthesis and virtual streamers. <xref ref-type="sec" rid="s8">Section 8</xref> discusses the ethical implications and reviews the progress and foundational challenges of generative AI. <xref ref-type="sec" rid="s9">Section 9</xref> concludes the paper by synthesizing the key findings and proposing directions for further research.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The generative intelligence framework: a structural overview</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Preliminary Knowledge</title>
<sec id="s2_1">
<label>2.1</label>
<title>Generative AI in Anime Production</title>
<p>The anime industry is undergoing notable changes precipitated by the growing emergence of Generative Artificial Intelligence (GAI). The synergistic integration of Natural Language Processing (NLP), Computer Vision (CV), and cross-modal synthesis offers a <italic>potential</italic> pathway to automate certain conventional manual workflows. While these advanced models incur high computational overhead, their ability to reduce human effort and accelerate content iteration <italic>can represent</italic> a strategic shift towards more efficient and cost-effective production paradigms. Language Models (LMs) and Diffusion Models have emerged as pivotal instruments, providing the field with enhanced content comprehension and generative capabilities. This significant influence of AI in creative industries necessitates a judicious equilibrium between technological innovation and the preservation of human ingenuity. While AI automates repetitive tasks across numerous sectors, within creative domains like anime, it functions as a collaborative instrument, enabling novel creative avenues, optimizing workflows, and enhancing creative processes [<xref ref-type="bibr" rid="ref-15">15</xref>]. Nevertheless, maintaining the human element and inherent authenticity that define the output of creative industries remains paramount [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>This significant evolution is impacting various echelons of anime production, from the foundational literary and visual substrates to certain culminating deliverables. Specifically, the following critical domains are undergoing <bold>notable developments</bold>:</p>
<p><bold>Narrative and Graphic Novel Genesis:</bold> Literary works, serving as rich repositories of imaginative narratives, frequently constitute the genesis for anime adaptations, while graphic novels (manga) represent a sophisticated synthesis of visual and textual storytelling. AI&#x2019;s role encompasses facilitating script generation and potentially influencing narrative architectures.</p>
<p><bold>Illustrative and Keyframe Synthesis:</bold> Illustrations, conveying nuanced emotional expression and visual narratives, and key animation, defining pivotal movement frames, are critical constituents of the animation process. AI is being leveraged to expedite the colorization of anime line drawings [<xref ref-type="bibr" rid="ref-9">9</xref>], synthesize anime-stylized imagery (Yang) [<xref ref-type="bibr" rid="ref-10">10</xref>], and even contribute to character conceptualization (Tang and Chen) [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p><bold>Episodic and Interactive Media Expansion:</bold> Anime television series, a cornerstone of the industry and a primary revenue stream, and interactive media, notably role-playing games (RPGs), amplify the influence of anime narratives through immersive engagement. AI contributes to enhanced efficiency across various stages of animation production, including pre-production, asset creation, animation production, and post-production [<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>The confluence of sophisticated LMs, such as GPT and BERT, which exhibit aptitude in generating coherent scripts and dialogues, with advanced Diffusion Models, adept at synthesizing anime-stylized visuals, offers <italic>potential</italic> solutions to enduring challenges pertaining to creative efficiency, content personalization, and fiscal optimization. This synergistic integration aims to bridge textual narratives with visual content, contributing to a period of notable advancement in certain aspects of the anime creation paradigm. This integration is facilitating <italic>some aspects of</italic> industrial upgrading, fostering innovation, and augmenting productivity within <italic>parts of</italic> the digital creative industry (Wagan and Sidra) [<xref ref-type="bibr" rid="ref-11">11</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Diffusion Models</title>
<p>Diffusion models [<xref ref-type="bibr" rid="ref-3">3</xref>], a class of generative models, leverage the principle of stochastic reverse diffusion to synthesize images from latent noise. The core mechanistic paradigm involves a forward diffusion process, iteratively corrupting data with Gaussian noise, followed by a reverse denoising process, reconstructing the image from the noise distribution. In essence, these models learn to invert the progressive degradation of data structure, enabling the recovery of high-fidelity images. This process is illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Overview of diffusion models (DDPM, SGM, and Score SDE diffusion and denoising processes)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-2.tif"/>
</fig>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Mathematical Formalization of Diffusion Processes</title>
<p>Diffusion models operate on the principle of incrementally transforming data through a <bold>forward diffusion process</bold> and subsequently reversing this transformation via a <bold>denoising process</bold> to generate novel samples.</p>
<p>The <bold>forward process</bold> systematically introduces Gaussian noise into an initial image x<sub>0</sub> across <bold>T</bold> discrete timesteps, progressively corrupting it until it approximates a pure noise distribution x<sub>T</sub>. This controlled degradation is governed by a predetermined noise schedule.</p>
<p>Conversely, the reverse denoising process endeavors to iteratively reconstruct the original image from noise. This is achieved by training a neural network to predict the subtle noise component at each timestep, effectively learning to reverse the corruption introduced by the forward process. The core objective during training is to parameterize this denoising network, enabling it to accurately approximate the conditional probability distribution of a slightly less noisy image given its noisy counterpart.</p>
<p><bold>Generative sampling</bold> leverages this trained denoising network. It commences with a random sample drawn from a Gaussian noise distribution, analogous to x<sub>T</sub>. The network then iteratively refines this noisy input through successive denoising steps, gradually transforming the pure noise into a coherent, synthesized image x<sub>0</sub>.</p>
<p>The optimization of diffusion models primarily involves minimizing the Evidence Lower Bound (ELBO). This objective function quantifies the discrepancy between the forward diffusion process and the model&#x2019;s learned reverse denoising capabilities, effectively guiding the network to accurately reverse the noise corruption.</p>
<p>For readers seeking a comprehensive mathematical treatment of these processes, we refer to the foundational works cited in the original text [<xref ref-type="bibr" rid="ref-18">18</xref>&#x2013;<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Development History of Diffusion Models</title>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the historical development of diffusion models. The trajectory of diffusion models has been marked by distinct phases, each characterized by pivotal advancements that have collectively propelled their capabilities from theoretical constructs to increasingly powerful generative tools, influencing fields like anime generation.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Historical development of diffusion models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-3.tif"/>
</fig>
<p>The initial phase, ignited by the advent of Denoising Diffusion Probabilistic Models (DDPMs), focused on solidifying theoretical foundations [<xref ref-type="bibr" rid="ref-2">2</xref>]. Key innovations like Denoising Diffusion Implicit Models (DDIMs) significantly enhanced sampling efficiency, substantially reducing the steps required for high-fidelity generation [<xref ref-type="bibr" rid="ref-22">22</xref>]. Concurrently, Classifier Guidance and its successor, Classifier-Free Guidance, markedly improved conditional image synthesis, laying the groundwork for the sophisticated text-to-image models that followed [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. Theoretical explorations into discrete diffusion models, exemplified by Multinomial Diffusion and D3PM, also contributed to this foundational period [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
<p>Subsequent developments centered on diversifying applications and enhancing scalability. Techniques such as Latent Diffusion and VQ Diffusion were instrumental in applying diffusion models to large-scale datasets, making them amenable to real-world applications [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. This period also saw significant strides in sampling acceleration, with methods like PNDM and Analytic-DPM that further reduced computational overhead while maintaining generative quality [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]. The observed versatility of these models has extended their utility beyond mere image generation to tasks like semantic segmentation and advanced image editing.</p>
<p>The focus then shifted toward the creation of large-scale models, particularly in the text-to-image domain. Influential models like DALLE-2 and Imagen showcased impressive capabilities in synthesizing images from textual prompts, leveraging vast datasets [<xref ref-type="bibr" rid="ref-28">28</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>]. The release of open-source initiatives like Stable Diffusion, alongside accompanying massive datasets such as Laion-5B, has made these powerful generative tools widely accessible [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. This has led to significant and rapid adoption within research communities, among individual creators, and in agile development environments, fostering experimentation and innovation in AI-driven content generation.</p>
<p>This widespread adoption in research and creative exploration has rapidly led to a phase of deployed applications and domain expansion, manifesting primarily in specialized tools, academic prototypes, and niche creative workflows, rather than comprehensive industry-wide overhauls. Stable Diffusion, in particular, became a cornerstone for diverse applications [<xref ref-type="bibr" rid="ref-3">3</xref>], including image inpainting (e.g., Equilibrium Diffusion, Shadow Diffusion) [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>], image perception, 3D generation (e.g., DreamFusion, Magic3D) [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>], video generation (e.g., Latent Video Diffusion) [<xref ref-type="bibr" rid="ref-34">34</xref>], and medical imaging (e.g., MedSegDiff) [<xref ref-type="bibr" rid="ref-35">35</xref>]. This marked a crucial transition from academic research to practical utility.</p>
<p>The most recent phase emphasizes controllability and cross-domain innovation. Tools like ControlNet have enabled precise manipulation of generated images through explicit conditions (e.g., edge maps, depth maps), offering enhanced creative control. Advancements in text-to-3D generation (e.g., Point-E, DreamFusion), and text-driven video synthesis (e.g., Video Diffusion Models, Make-A-Video) further extended their capabilities [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>&#x2013;<xref ref-type="bibr" rid="ref-38">38</xref>]. Concurrently, ongoing research into computational efficiency (e.g., Latent Diffusion [<xref ref-type="bibr" rid="ref-3">3</xref>], Efficient Diffusion) and multi-modal fusion (e.g., Diffusion-LM) continues to explore and enhance the capabilities of diffusion models, integrating disparate data types for more complex and nuanced generative tasks, including their observed impact on anime generation through specialized applications and stylistic control [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Datasets for Diffusion Model Training</title>
<p>The training of diffusion models, irrespective of modality (image, video, audio), necessitates large-scale, high-fidelity datasets [<xref ref-type="bibr" rid="ref-41">41</xref>]. Optimal datasets are characterized by: <bold>Extensive Cardinality:</bold> Datasets comprising hundreds of millions to billions of paired data samples (e.g., image-text) are requisite for capturing the inherent diversity and complexity of the data manifold [<xref ref-type="bibr" rid="ref-42">42</xref>]. <bold>Comprehensive Heterogeneity:</bold> Datasets must encompass a broad spectrum of scenes, styles, and linguistic representations to ensure robust generalization of generated outputs [<xref ref-type="bibr" rid="ref-43">43</xref>]. <bold>Precise Annotation Fidelity:</bold> Particularly for conditional generative tasks, the semantic coherence between textual and visual/temporal data is paramount. Exemplary datasets are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Datasets for diffusion models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Dataset name</th>
<th>Release date</th>
<th>Sample size</th>
<th>Primary use</th>
<th>Characteristics</th>
</tr>
</thead>
<tbody>
<tr>
<td>LAION-5B [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>2022-03</td>
<td>5.85 Billion image-text Pairs</td>
<td>Text-to-image generation, multilingual research</td>
<td>Multilingual, large-scale</td>
</tr>
<tr>
<td>Re-LAION-5B [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>2024-08</td>
<td>5.85 Billion (Updated)</td>
<td>Safety research, language-visual learning</td>
<td>Enhanced safety</td>
</tr>
<tr>
<td>ImageNet [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td>2009</td>
<td>1.2 Million images</td>
<td>Unconditional image generation, classification</td>
<td>Hierarchical classification structure</td>
</tr>
<tr>
<td>COCO [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td>2017</td>
<td>120,000 images</td>
<td>Scene generation, object detection</td>
<td>Rich annotations</td>
</tr>
<tr>
<td>CelebA [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td>2015</td>
<td>203,000 Images</td>
<td>Face generation, attribute recognition</td>
<td>40 facial attribute annotations, aligned faces</td>
</tr>
<tr>
<td>WebVid-10M [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>2021-03</td>
<td>10.7 Million video clips</td>
<td>Text-to-video generation</td>
<td>High diversity</td>
</tr>
<tr>
<td>HD-Vila-100M [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>2022</td>
<td>100 Million video clips</td>
<td>General video generation</td>
<td>High definition, large-scale</td>
</tr>
<tr>
<td>VidProM [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>2024-03</td>
<td>6.69 Million videos</td>
<td>Prompt evaluation, model comparison</td>
<td>Novel, evaluation-oriented</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Exemplary Datasets for Diffusion Model Training:</p>
<p><bold>LAION-5B (March 2022) [<xref ref-type="bibr" rid="ref-12">12</xref>]:</bold> A corpus of 5.85 billion image-text pairs, curated via CLIP-based filtering, exhibiting multilingual scope. This dataset has facilitated the training of models such as Stable Diffusion, enhancing generative fidelity and zero-shot capabilities.</p>
<p><bold>Re-LAION-5B (August 2024)</bold> <bold>[<xref ref-type="bibr" rid="ref-44">44</xref>]:</bold> An augmented iteration of LAION-5B, incorporating stringent filtering to mitigate illicit content and providing academically compliant subsets, thereby addressing ethical and legal considerations.</p>
<p><bold>WebVid-10M (March 2021) [<xref ref-type="bibr" rid="ref-45">45</xref>]:</bold> A video dataset comprising 10.7 million video clips (52,000 h) with alt-text annotations, contributing to improved temporal consistency and zero-shot video generation.</p>
<p><bold>HD-Vila-100M (2022) [<xref ref-type="bibr" rid="ref-46">46</xref>]:</bold> A large-scale video dataset consisting of 100 million high-definition videos (371,000 h) with automatically transcribed textual data, supporting generalized video synthesis tasks.</p>
<p><bold>VidProM (March 2024) [<xref ref-type="bibr" rid="ref-47">47</xref>]:</bold> A synthetic video dataset featuring 6.69 million generated video clips (1.6-3 s each), synthesized using multiple generative models, and augmented with NSFW detection and prompt embeddings, designed to facilitate model evaluation and prompt engineering research.</p>
</sec>
<sec id="s2_2_4">
<label>2.2.4</label>
<title>Architectural Advantages</title>
<p>The architectural design of Diffusion Models offers several advantages in generative tasks, notably in latent space operation, flexibility and tractability, neural network adaptability, conditional generation capabilities, and scalability. Many Diffusion Models, such as Stable Diffusion, employ a two-stage training paradigm that first compresses high-dimensional image data into a lower-dimensional latent space via an autoencoder [<xref ref-type="bibr" rid="ref-44">44</xref>]. This approach not only substantially reduces computational complexity but also enhances model scalability, particularly for high-resolution image synthesis. By conducting diffusion and reverse diffusion within this latent space, models like Latent Diffusion Models (LDMs) can efficiently process intricate data while often preserving high-quality generative outcomes [<xref ref-type="bibr" rid="ref-3">3</xref>]. Diffusion Models can effectively balance the analytical tractability of simpler distributions (e.g., Gaussian) with the expressive power of complex models (e.g., GANs), enabling them to model sophisticated data distributions with notable training stability and sampling efficiency, and have often outperformed traditional generative models as evidenced by their superiority over GANs in image synthesis [<xref ref-type="bibr" rid="ref-51">51</xref>]. The reverse diffusion process is typically orchestrated by flexible neural network architectures, including U-Net or Transformer variants, allowing for task-specific customization; for instance, Stable Diffusion leverages U-Net with cross-attention, while Stable Diffusion 3 employs a Diffusion Transformer (DiT), showcasing this architectural versatility [<xref ref-type="bibr" rid="ref-52">52</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>]. Furthermore, Diffusion Models excel in conditional generation through mechanisms like cross-attention or dedicated conditioning modules (e.g., text encoders), enabling models such as Stable Diffusion and DALL-E 2 to synthesize imagery from textual prompts, thereby broadening their applicability to tasks like text-to-image and layout-to-image generation [<xref ref-type="bibr" rid="ref-28">28</xref>]. Their inherent scalability, facilitated by latent space representations and efficient architectures, allows them to manage high-dimensional data, balancing generative quality with computational expediency; Cascade Diffusion Models, for example, have demonstrated enhanced high-resolution image generation through multi-stage diffusion processes [<xref ref-type="bibr" rid="ref-54">54</xref>].</p>
</sec>
<sec id="s2_2_5">
<label>2.2.5</label>
<title>Training Methodology</title>
<p>The training regimen for Diffusion Models encompasses several critical steps, primarily involving the forward and reverse diffusion processes, the strategic selection of variance schedules, and the application of advanced optimization techniques. The training commences with the forward diffusion process, wherein original data is progressively corrupted by Gaussian noise, transitioning from the data distribution to a pure noise distribution. This process, governed by a predefined variance schedule (typically linear or cosine), functions as a Markov chain, incrementally introducing noise. The judicious choice of variance scheduling significantly impacts training stability and generative quality. Subsequently, the reverse diffusion process trains a neural network to invert this corruption, iteratively denoising from pure noise to reconstruct the original data. The neural network typically predicts the noise or the denoised data at each step, optimizing through the minimization of a simple mean squared error (MSE) based loss function. Variance schedules, which dictate the rate of noise addition, are paramount, with linear and cosine schedules being common choices. Advanced training techniques, such as DDIM (Denoising Diffusion Implicit Models), have markedly accelerated sampling by reducing the number of necessary steps. Progressive Distillation further refines model performance using a teacher-student framework, and Consistency Models enhance generative quality by enforcing specific consistency constraints [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Language Model</title>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>Foundational Architectures of Transformer-Based Language Models</title>
<p>The historical development of these models is summarized in <xref ref-type="table" rid="table-2">Table 2</xref>. Transformer-based language models, fundamental to advancements in natural language processing, are broadly categorized into three core architectural paradigms, each optimized for distinct linguistic tasks [<xref ref-type="bibr" rid="ref-64">64</xref>]:</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Historical development of large language models</title>
</caption>
<table>
<colgroup>
<col/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Year</th>
<th align="left">Model name</th>
<th align="left">Open/closed source</th>
<th align="left">Model type</th>
<th align="left">Dev. Team/Co.</th>
</tr>
</thead>
<tbody>
<tr>
<td>2018</td>
<td>BERT [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td>Open source</td>
<td>Encoder</td>
<td>Google</td>
</tr>
<tr>
<td>2019</td>
<td>GPT-2 [<xref ref-type="bibr" rid="ref-55">55</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2020</td>
<td>RoBERTa [<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
<td>Open source</td>
<td>Encoder</td>
<td>Facebook</td>
</tr>
<tr>
<td>2020</td>
<td>T5 [<xref ref-type="bibr" rid="ref-57">57</xref>]</td>
<td>Open source</td>
<td>Enc-Dec</td>
<td>Google</td>
</tr>
<tr>
<td>2020</td>
<td>GPT-3 [<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2021</td>
<td>GPT-Neo [<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
<td>Open source</td>
<td>Decoder</td>
<td>EleutherAI</td>
</tr>
<tr>
<td>2022</td>
<td>LaMDA [<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>Google</td>
</tr>
<tr>
<td>2022</td>
<td>PaLM [<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>Google</td>
</tr>
<tr>
<td>2022</td>
<td>GPT-3 [<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2023</td>
<td>GPT-4 [<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2023</td>
<td>LLaMA [<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
<td>Open source</td>
<td>Decoder</td>
<td>Meta</td>
</tr>
<tr>
<td>2023</td>
<td>Bard</td>
<td>Closed source</td>
<td>Decoder</td>
<td>Google</td>
</tr>
<tr>
<td>2023</td>
<td>Claude</td>
<td>Closed source</td>
<td>Decoder</td>
<td>Anthropic</td>
</tr>
<tr>
<td>2024</td>
<td>Gemma 2</td>
<td>Open source</td>
<td>Decoder</td>
<td>Google</td>
</tr>
<tr>
<td>2024</td>
<td>Llama 3.1</td>
<td>Open source</td>
<td>Decoder</td>
<td>Meta</td>
</tr>
<tr>
<td>2024</td>
<td>OpenAI o1</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2024</td>
<td>Phi 3.5</td>
<td>Open source</td>
<td>Decoder</td>
<td>Microsoft</td>
</tr>
<tr>
<td>2024</td>
<td>Qwen2.5</td>
<td>Open source</td>
<td>Decoder</td>
<td>Alibaba Cloud</td>
</tr>
<tr>
<td>2025</td>
<td>GPT-4.5</td>
<td>Closed source</td>
<td>Decoder</td>
<td>OpenAI</td>
</tr>
<tr>
<td>2025</td>
<td>DeepSeek R1</td>
<td>Open source</td>
<td>Decoder</td>
<td>DeepSeek</td>
</tr>
<tr>
<td>2025</td>
<td>Gemini 2.0</td>
<td>Closed source</td>
<td>Decoder</td>
<td>GoogleDeepMind</td>
</tr>
<tr>
<td>2025</td>
<td>Grok-3</td>
<td>Closed source</td>
<td>Decoder</td>
<td>xAI</td>
</tr>
<tr>
<td>2025</td>
<td>Claude 3.7 Sonnet</td>
<td>Closed source</td>
<td>Decoder</td>
<td>Anthropic</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Encoder-Centric Models:</bold> Exemplified by models such as BERT, these architectures are primarily designed for comprehension-oriented tasks [<xref ref-type="bibr" rid="ref-7">7</xref>]. They excel at understanding the intricate semantic relationships within input text by learning deep contextual representations. This makes them highly effective for applications like text classification, named entity recognition, and reading comprehension, where the goal is to extract meaning from existing text.</p>
<p><bold>Encoder-Decoder Hybrid Models:</bold> Represented by models like T5, these architectures combine the strengths of both encoders and decoders [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. The encoder processes the input sequence, capturing its semantic essence, while the decoder then uses this understanding to synthesize an output sequence. This dual mechanism makes them highly adept at sequence transduction tasks, such as machine translation and text summarization, where one sequence is transformed into another.</p>
<p><bold>Decoder-Dominant Models:</bold> Typified by the GPT series (e.g., GPT-4, LLaMA), these models are built for generative tasks [<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. Operating autoregressively, they predict subsequent tokens based on preceding ones, enabling them to produce fluent and contextually coherent text. Their inherent design makes them ideal for applications requiring creative text generation, dialogue systems, and content creation, including potential applications in anime script and narrative generation.</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Encoder-Centric Models: Contextual Representation Learning</title>
<p><bold>Encoder-centric models</bold>, notably <bold>BERT</bold> [<xref ref-type="bibr" rid="ref-7">7</xref>], are specifically engineered to grasp contextual information within text. Their pre-training focuses on learning deep relationships between words in a given context, making them highly effective at tasks that require understanding existing text, such as text classification or question-answering. However, their architecture, being primarily focused on comprehension, inherently limits their direct application in generating novel text or dialogues.</p>
</sec>
<sec id="s2_3_3">
<label>2.3.3</label>
<title>Encoder-Decoder Hybrid Models: Sequence Transduction</title>
<p>The architecture of an encoder-decoder model is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. Encoder-decoder models facilitate sequence transduction by employing a two-part system. An encoder first processes an input sequence to distill its underlying meaning into a rich contextual representation. This representation is then passed to a decoder, which uses this understanding to construct a new, coherent output sequence. This architecture is highly effective for tasks where the input needs to be transformed into a different output format, such as translating a script from one language to another or summarizing a long narrative into a concise synopsis, which could be valuable for managing anime production content [<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>].</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Encoder-decoder model architecture diagram</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-4.tif"/>
</fig>
</sec>
<sec id="s2_3_4">
<label>2.3.4</label>
<title>Decoder-Dominant Models: Autoregressive Sequence Generation</title>
<p>The schematic for a decoder-only model is depicted in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. Decoder-dominant models operate autoregressively, generating sequences token-by-token, conditioned on preceding tokens. Input sequences are embedded and processed through stacked decoder blocks, each comprising masked self-attention, add-and-norm, and FFN layers. The output is a probability distribution over subsequent tokens.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Decoder-only model architecture schematic</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-5.tif"/>
</fig>
<p>The GPT and LLaMA series are representative decoder-dominant models [<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. GPT models demonstrate strong generative capabilities, with GPT-4 extending to multimodal content. LLaMA models offer scalable architectures for diverse applications.</p>
</sec>
<sec id="s2_3_5">
<label>2.3.5</label>
<title>Architectural Advantages</title>
<p>The Transformer architecture derives its notable generative capabilities from several key innovations, including the self-attention mechanism, inherent parallelizability, scalability, the absence of recurrent structures, multi-head attention, positional encodings, and a continuous evolution of efficient variants. The self-attention mechanism, a cornerstone of the Transformer, dynamically allocates attention weights based on the relevance of input sequence elements, thereby capturing long-range dependencies without relying on sequential processing, a significant advancement over traditional models in complex tasks like machine translation and text generation [<xref ref-type="bibr" rid="ref-64">64</xref>]. Unlike recurrent neural networks (RNNs), Transformers inherently support parallel processing of entire input sequences, dramatically accelerating training and inference, particularly on GPUs [<xref ref-type="bibr" rid="ref-64">64</xref>]. This parallelization, combined with the ability to handle arbitrary sequence lengths, provides Transformers with high scalability across diverse tasks, enabling models such as BERT and GPT to adapt from text classification to generation via pre-training and fine-tuning [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-66">66</xref>]. The elimination of recurrent structures mitigates the vanishing gradient problem, facilitating the training of deeper, more stable, and efficient networks for long sequences. Multi-head attention further augments representational capacity by allowing the model to concurrently attend to distinct subspaces within the input sequence, instrumental in capturing bidirectional context (e.g., BERT) or facilitating autoregressive generation (e.g., GPT) [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-66">66</xref>]. Given the inherent order-agnostic nature of self-attention, Transformers incorporate positional encodings, such as sinusoidal functions, to convey token position; recent innovations like Rotary Positional Embeddings (RoPE) and ALiBi optimize for relative positional dependencies and enable fine-tuning on longer sequences after pre-training on shorter ones [<xref ref-type="bibr" rid="ref-67">67</xref>,<xref ref-type="bibr" rid="ref-68">68</xref>]. To address the computational demands of the Transformer, numerous efficient variants have emerged, including Reformer, which leverages locality-sensitive hashing (LSH) to reduce attention complexity from O(N<sup>2</sup>) to O(NlogN), and BigBird, achieving O(N) complexity via small-world networks [<xref ref-type="bibr" rid="ref-69">69</xref>,<xref ref-type="bibr" rid="ref-70">70</xref>]. Furthermore, FlashAttention and FlashAttention-2 have dramatically accelerated attention computations, reaching speeds up to 230 TFLOPs/s on A100 GPUs [<xref ref-type="bibr" rid="ref-71">71</xref>]. Contemporary trends in Transformer architecture, as of 2025, focus on sparsity through sparse attention mechanisms for reduced computation and improved memory efficiency, Mixture-of-Experts (MoE) models to enhance scalability and efficiency by partitioning the network into specialized modules, and adaptive computation techniques that dynamically adjust computational resources based on input complexity to optimize performance.</p>
</sec>
<sec id="s2_3_6">
<label>2.3.6</label>
<title>Training Methodology</title>
<p>Transformer training typically unfolds in a two-stage paradigm: pre-training and fine-tuning. Initial pre-training occurs on vast corpora (e.g., Wikipedia) using self-supervised objectives such as masked language modeling (BERT) or auto-regressive language modeling (GPT). This is followed by fine-tuning on smaller, task-specific datasets to adapt the model for particular applications, a methodology that significantly curtails training costs and enhances generalization. The standard Transformer architecture comprises an encoder and a decoder. The encoder processes input sequences via multi-head self-attention to generate contextual representations, while the decoder, employing masked self-attention and encoder-decoder attention, produces output sequences auto-regressively. Attention mechanisms are pivotal in training: encoder self-attention captures intra-input relationships; decoder masked self-attention ensures predictions are solely based on preceding tokens; and encoder-decoder attention integrates the encoder&#x2019;s output into the decoder, thereby enhancing generative quality. Layer Normalization and Residual Connections, applied after each sub-layer, are crucial for mitigating the vanishing gradient problem and facilitating the training of deeper networks. Optimization for Transformer training typically employs the Adam optimizer coupled with learning rate schedules (e.g., warmup and decay) to improve convergence, while regularization techniques like Dropout prevent overfitting, especially in models with extensive parameter counts. Recent advancements in training methodology include efficient inference techniques such as Key-Value Caching to obviate redundant computations of key and value vectors, speculative decoding, and multi-token prediction to balance accuracy and speed in real-time applications. Pre-layer normalization (Pre-LN) has been introduced to enhance training stability. Moreover, the Transformer architecture has been extended for multimodal training, exemplified by models like DALL-E, which jointly process complex datasets encompassing both text and images [<xref ref-type="bibr" rid="ref-64">64</xref>].</p>
</sec>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Model Fine-Tuning Techniques and Anime-Related Datasets</title>
<sec id="s2_4_1">
<label>2.4.1</label>
<title>Model Fine-Tuning Techniques</title>
<p>An overview of common fine-tuning techniques is provided in <xref ref-type="table" rid="table-3">Table 3</xref>. Full Fine-Tuning, the most direct method, updates all model parameters. While offering maximal adaptation potential for a specific style, it incurs high computational expense, requires substantial data, and exhibits limited generalizability to diverse, unseen anime styles, risking overfitting on limited datasets [<xref ref-type="bibr" rid="ref-72">72</xref>].</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Overview of common model fine-tuning techniques</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Technique name</th>
<th align="left">Primary model type</th>
<th align="left">Description</th>
<th align="left">Modifies original model weights?</th>
<th align="left">Parameter efficiency</th>
</tr>
</thead>
<tbody>
<tr>
<td>LoRA (Low-Rank Adaptation)</td>
<td>LLMs, diffusion models</td>
<td>Freezes pre-trained model weights, injects trainable low-rank decomposition matrices, significantly reducing trainable parameters.</td>
<td>No</td>
<td>High</td>
</tr>
<tr>
<td>ControlNet</td>
<td>Diffusion models</td>
<td>Adds additional conditioning (e.g., edges, poses) to diffusion models to control image generation, providing precise guidance.</td>
<td>No</td>
<td>Medium</td>
</tr>
<tr>
<td>DreamBooth</td>
<td>Diffusion models</td>
<td>Fine-tunes a diffusion model using a small number of images of a specific subject to generate new images of that subject in different scenarios.</td>
<td>Yes</td>
<td>Low</td>
</tr>
<tr>
<td>Textual inversion</td>
<td>Diffusion models</td>
<td>Learns new embeddings for specific visual concepts, enabling diffusion models to generate images based on these concepts.</td>
<td>No</td>
<td>High</td>
</tr>
<tr>
<td>HyperNetworks</td>
<td>Diffusion models, LLMs</td>
<td>A neural network generates the weights of another neural network, enabling efficient and flexible model adaptation.</td>
<td>Yes</td>
<td>High (Potential)</td>
</tr>
<tr>
<td>Full fine-tuning</td>
<td>LLMs, diffusion models</td>
<td>Trains all parameters of a pre-trained model on a new dataset to adapt to a specific task, incurring high computational cost.</td>
<td>Yes</td>
<td>Low</td>
</tr>
<tr>
<td>Feature-based fine-tuning</td>
<td>LLMs</td>
<td>Fine-tunes only the later layers of a pre-trained model, keeping earlier layers frozen, preserving general features for task adaptation.</td>
<td>Partial</td>
<td>Medium</td>
</tr>
<tr>
<td>Adapters</td>
<td>LLMs</td>
<td>Adds small task-specific modules to large language models, enabling parameter-efficient fine-tuning.</td>
<td>No</td>
<td>High</td>
</tr>
<tr>
<td>Prompt tuning</td>
<td>LLMs</td>
<td>Guides large language model behavior by adjusting input prompts without modifying model weights, offering lightweight adaptation.</td>
<td>No</td>
<td>Extremely high</td>
</tr>
<tr>
<td>Prefix-tuning</td>
<td>LLMs</td>
<td>Adds small trainable modules (prefixes) before the input sequence, enabling task-specific adaptation while keeping the model frozen.</td>
<td>No</td>
<td>High</td>
</tr>
<tr>
<td>Representation Fine-Tuning (ReFT)</td>
<td>LLMs</td>
<td>Learns task-specific interventions on hidden representations instead of updating weights, achieving parameter-efficient fine-tuning.</td>
<td>No</td>
<td>Extremely high</td>
</tr>
<tr>
<td>Instruction fine-tuning</td>
<td>LLMs</td>
<td>Fine-tunes large language models using instruction-response pairs to enhance their ability to follow instructions and generate task-specific responses.</td>
<td>Yes</td>
<td>Medium</td>
</tr>
<tr>
<td>Multi-task fine-tuning</td>
<td>LLMs</td>
<td>Simultaneously fine-tunes a single model on multiple related tasks, learning shared representations and improving performance across tasks.</td>
<td>Yes</td>
<td>Medium</td>
</tr>
<tr>
<td>Domain-specific fine-tuning</td>
<td>LLMs, diffusion models</td>
<td>Further trains a pre-trained model on a dataset specific to a particular domain, enhancing performance on tasks within that domain.</td>
<td>Yes</td>
<td>Medium</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To mitigate these costs, Parameter-Efficient Fine-Tuning (PEFT) methods update only a small subset of parameters or introduce minimal new ones. LoRA injects low-rank matrices, achieving high efficiency but struggling with blending or switching between vastly different styles [<xref ref-type="bibr" rid="ref-73">73</xref>]. Adapters insert small modules, offering modularity and efficiency at the cost of potential inference latency and configuration complexity; combining style-specific Adapters for generalization is challenging [<xref ref-type="bibr" rid="ref-74">74</xref>]. Prompt Tuning and Prefix-Tuning manipulate input embeddings for extreme efficiency but provide limited control over complex stylistic details [<xref ref-type="bibr" rid="ref-75">75</xref>,<xref ref-type="bibr" rid="ref-76">76</xref>]. Representation Fine-Tuning (ReFT) intervenes on latent representations, offering low cost and non-invasiveness, though its efficacy for complex anime style generalization is under investigation [<xref ref-type="bibr" rid="ref-77">77</xref>].</p>
<p>Diffusion-specific techniques provide conditional control. ControlNet adds spatial conditioning (e.g., edges), potent for structure but relying on the base model&#x2019;s style aptitude [<xref ref-type="bibr" rid="ref-78">78</xref>]. DreamBooth specializes models on specific subjects from few images, achieving high fidelity but overfitting and limiting subject generalization across varied anime styles [<xref ref-type="bibr" rid="ref-79">79</xref>]. Textual Inversion learns new embeddings for concepts, offering efficiency but limited capacity for intricate style representation [<xref ref-type="bibr" rid="ref-80">80</xref>]. HyperNetworks dynamically generate model parameters, promising flexibility but facing training instability and potentially yielding lower quality outputs than direct fine-tuning [<xref ref-type="bibr" rid="ref-81">81</xref>]. Feature-Based Fine-Tuning updates only later layers, computationally lighter but potentially insufficient for style changes requiring lower-level feature modification and bounded in its generalization to diverse styles [<xref ref-type="bibr" rid="ref-82">82</xref>].</p>
<p>Finally, Instruction Fine-Tuning, Multi-Task Fine-Tuning, and Domain-Specific Fine-Tuning enhance task-specific or domain-confined performance. While improving utility within a prescribed scope (e.g., sci-fi anime), they face constraints in achieving comprehensive generalization or synthesizing novel anime styles due to dataset limitations, task interference, and inherent model capacity [<xref ref-type="bibr" rid="ref-83">83</xref>].</p>
<p>While fine-tuning is crucial for anime style acquisition, achieving seamless, high-fidelity generalization across the diverse spectrum of anime aesthetics remains a significant research challenge.</p>
</sec>
<sec id="s2_4_2">
<label>2.4.2</label>
<title>Anime-Related Datasets</title>
<p><xref ref-type="table" rid="table-4">Table 4</xref> provides an overview of key anime-related datasets. Despite advancements, current anime/manga datasets face notable limitations impeding sophisticated model development. Prominent among these is pervasive stylistic heterogeneity without granular annotation, hindering style-specific mastery or seamless generalization across diverse aesthetics [<xref ref-type="bibr" rid="ref-99">99</xref>,<xref ref-type="bibr" rid="ref-100">100</xref>]. This often results in dataset bias, overrepresenting popular styles and diminishing performance on less common ones, exacerbated by long-tailed character distributions.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Overview of key anime-related datasets</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Year</th>
<th align="left">Conference /Journal</th>
<th align="left">Title</th>
<th align="left">Scale</th>
</tr>
</thead>
<tbody>
<tr>
<td>2025</td>
<td>COLING</td>
<td>Context-informed machine translation of manga using multimodal large language models [<xref ref-type="bibr" rid="ref-84">84</xref>]</td>
<td>Constructed a novel evaluation corpus, representing the premiere Japanese-Polish parallel corpus specifically curated for manga translation.</td>
</tr>
<tr>
<td>2024</td>
<td>Arxiv</td>
<td>Tails Tell Tales: chapter-wide manga transcriptions with character names [<xref ref-type="bibr" rid="ref-85">85</xref>]</td>
<td>Augments the extant PopManga dataset and established a novel character repository featuring over 11,000 distinct characters.</td>
</tr>
<tr>
<td>2024</td>
<td>Arxiv</td>
<td>CoMix: a comprehensive benchmark for multi-task comic understanding [<xref ref-type="bibr" rid="ref-86">86</xref>]</td>
<td>Comprises 3800 images sampled from 100 distinct manga volumes, meticulously annotated with 130,000 object instances, 30,000 text-character linkages, and 33,000 character clusters, encompassing 16,000 unique names.</td>
</tr>
<tr>
<td>2023</td>
<td>CVPR</td>
<td>Human-Art: a versatile human-centric dataset bridging natural and artificial scenes [<xref ref-type="bibr" rid="ref-87">87</xref>]</td>
<td>Features 50,000 high-fidelity images, exceeding 123,000 human instances, and spanning 20 diverse scene categories.</td>
</tr>
<tr>
<td>2023</td>
<td>Arxiv</td>
<td>Manga109Dialog: a large-scale dialogue dataset for comics speaker detection [<xref ref-type="bibr" rid="ref-88">88</xref>]</td>
<td>Comprises 132,692 speaker-text pairs and 9904 images featuring speaker-text pair annotations. Represents the world&#x2019;s largest dialogue corpus specifically curated for manga speaker detection.</td>
</tr>
<tr>
<td>2023</td>
<td>ACM-TG</td>
<td>Semi-supervised reference-based sketch extraction using a contrastive learning framework [<xref ref-type="bibr" rid="ref-89">89</xref>]</td>
<td>Curated a novel dataset of authentic sketches rendered by professional artists, encompassing four distinct sketching styles, each style paired with 25 corresponding color images.</td>
</tr>
<tr>
<td>2023</td>
<td>ACM-TG</td>
<td>Parsing-Conditioned Anime Translation: A New Dataset and Method [<xref ref-type="bibr" rid="ref-90">90</xref>]</td>
<td>Introduced the Danbooru-Parsing dataset, featuring 4921 images with dense annotations.</td>
</tr>
<tr>
<td>2022</td>
<td>NeurIPS-DB</td>
<td>AnimeRun: 2D animation visual correspondence from open source 3D Movies [<xref ref-type="bibr" rid="ref-91">91</xref>]</td>
<td>A novel corpus for establishing 2D animation visual correspondence.</td>
</tr>
<tr>
<td>2022</td>
<td>ECCV</td>
<td>COO: comic onomatopoeia dataset for recognizing arbitrary or truncated texts [<xref ref-type="bibr" rid="ref-92">92</xref>]</td>
<td>Comprises 10,602 images derived from 109 manga volumes, featuring 61,465 polygon annotations and 2261 linkages.</td>
</tr>
<tr>
<td>2022</td>
<td>ECCV</td>
<td>AnimeCeleb: large-scale animation celebheads dataset for head reenactment [<xref ref-type="bibr" rid="ref-93">93</xref>]</td>
<td>A substantial-scale corpus of animation character heads specifically designed for head reenactment tasks.</td>
</tr>
<tr>
<td>2021</td>
<td>Arxiv</td>
<td>DAF:RE: a challenging, crowd-sourced, large-scale, long-tailed dataseT For Anime Character Recognition [<xref ref-type="bibr" rid="ref-94">94</xref>]</td>
<td>Comprises nearly 500,000 images, spanning over 3000 distinct categories. Represents a large-scale, crowd-sourced, long-tailed benchmark dataset.</td>
</tr>
<tr>
<td>2020</td>
<td>ACM-MM</td>
<td>Cartoon face recognition: a benchmark dataset [<xref ref-type="bibr" rid="ref-95">95</xref>]</td>
<td>The iCartoonFace dataset comprises 389,678 images and features 5013 distinct cartoon characters for recognition tasks; it also includes 60,000 images with annotations for 109,810 cartoon faces for detection tasks. At the time of its release, it constituted the largest benchmark dataset for cartoon face recognition.</td>
</tr>
<tr>
<td>2020</td>
<td>ECCV Workshop</td>
<td>Unconstrained text detection in manga: a new dataset and baseline [<xref ref-type="bibr" rid="ref-96">96</xref>]</td>
<td>Curated a dataset encompassing 450 images accompanied by text segmentation annotations, sourced from the Manga109 dataset.</td>
</tr>
<tr>
<td>2020</td>
<td>ECCV</td>
<td>DanbooRegion: an illustration region dataset [<xref ref-type="bibr" rid="ref-97">97</xref>]</td>
<td>Features 5377 pairs of region annotations, capturing diverse artistic region compositions.</td>
</tr>
<tr>
<td>2020</td>
<td>MMUL</td>
<td>Building a manga dataset &#x201C;Manga109&#x201D; with annotations for multimedia applications [<xref ref-type="bibr" rid="ref-98">98</xref>]</td>
<td>Comprises 109 Japanese manga volumes authored by 94 creators, totaling 21,142 pages, with over 500,000 annotations encompassing panels, text boxes, faces, and body poses. Upon its introduction, it represented the most extensive publicly available manga image dataset for research purposes.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Furthermore, the landscape is characterized by fragmentation, with datasets often task-specific and lacking comprehensive multi-modal integration of visual, textual, and structural elements crucial for holistic understanding [<xref ref-type="bibr" rid="ref-101">101</xref>]. Annotation quality and consistency remain concerns, prone to human error and complexity, particularly for intricate tasks [<xref ref-type="bibr" rid="ref-102">102</xref>]. Finally, the static nature of most datasets fails to capture the domain&#x2019;s dynamic evolution, limiting their enduring relevance for cutting-edge research. These constraints collectively underscore the pressing need for more nuanced, integrated, and continuously evolving data resources.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Image Generation</title>
<p>This section elucidates the current state of generative image synthesis models, focusing on prominent architectures particularly relevant to anime content creation, a domain where artificial intelligence is increasingly employed as an artistic tool. Our selection encompasses architectures distinguished by their efficacy in generating high-fidelity anime-style visuals, their impact on the field, and the depth of available research and practical applications. We focus on Stable Diffusion [<xref ref-type="bibr" rid="ref-3">3</xref>], related mainstream models, manga synthesis methodologies, and evaluation metrics.</p>
<p>While the substantial parameters, training compute, and VRAM requirements detailed in <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> underscore the significant computational investment in these models, these costs are often offset by significant gains in creative efficiency and reduced manual labor. Moreover, ongoing research into architectural efficiencies and optimization techniques actively seeks to enhance their accessibility and integration into practical creative pipelines, continuously lowering the effective cost of deployment relative to the benefits derived.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Mainstream image generation large models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model Name</th>
<th align="left">Rel. Date</th>
<th align="left">Dev. Team</th>
<th align="left">Params</th>
<th align="left">Open Source</th>
</tr>
</thead>
<tbody>
<tr>
<td>Stable diffusion 1.5 [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Oct 2022</td>
<td>StabilityAI</td>
<td>983M</td>
<td>Yes</td>
</tr>
<tr>
<td>Stable diffusion XL [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Jul 2023</td>
<td>StabilityAI</td>
<td>3.5B</td>
<td>Yes</td>
</tr>
<tr>
<td>Stable diffusion 3.0 [<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>Feb 2024</td>
<td>StabilityAI</td>
<td>800M-8B</td>
<td>Yes</td>
</tr>
<tr>
<td>Stable diffusion 3.5 [<xref ref-type="bibr" rid="ref-53">53</xref>]</td>
<td>Oct 2024</td>
<td>StabilityAI</td>
<td>2.5B-8B</td>
<td>Yes</td>
</tr>
<tr>
<td>FLUX.1 [<xref ref-type="bibr" rid="ref-103">103</xref>]</td>
<td>Aug 2024</td>
<td>Black forest labs</td>
<td>12B</td>
<td>Yes</td>
</tr>
<tr>
<td>Midjourney V7 [<xref ref-type="bibr" rid="ref-104">104</xref>]</td>
<td>Apr 2025</td>
<td>Midjourney</td>
<td>Undisc.</td>
<td>No</td>
</tr>
<tr>
<td>DALL&#x00B7;E 3 [<xref ref-type="bibr" rid="ref-105">105</xref>]</td>
<td>Aug 2023</td>
<td>OpenAI</td>
<td>Undisc.</td>
<td>No</td>
</tr>
<tr>
<td>Imagen 3 [<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>Aug 2024</td>
<td>Google</td>
<td>Undisc.</td>
<td>No</td>
</tr>
<tr>
<td>Firefly [<xref ref-type="bibr" rid="ref-106">106</xref>]</td>
<td>Mar 2023</td>
<td>Adobe</td>
<td>Undisc.</td>
<td>No</td>
</tr>
<tr>
<td>Ideogram 3.0 [<xref ref-type="bibr" rid="ref-107">107</xref>]</td>
<td>Mar 2025</td>
<td>Ideogram AI</td>
<td>4B</td>
<td>No</td>
</tr>
<tr>
<td>CogView-4 [<xref ref-type="bibr" rid="ref-108">108</xref>]</td>
<td>Mar 2025</td>
<td>Zhipu AI</td>
<td>6B</td>
<td>Yes</td>
</tr>
<tr>
<td>Kolors [<xref ref-type="bibr" rid="ref-109">109</xref>]</td>
<td>May 2024</td>
<td>Kuaishou</td>
<td>3B</td>
<td>Yes</td>
</tr>
<tr>
<td>Seedream 3.0 [<xref ref-type="bibr" rid="ref-110">110</xref>]</td>
<td>Apr 2025</td>
<td>ByteDance</td>
<td>Undisc.</td>
<td>No</td>
</tr>
<tr>
<td>HunyuanDiT [<xref ref-type="bibr" rid="ref-111">111</xref>]</td>
<td>May 2024</td>
<td>Tencent</td>
<td>15B</td>
<td>Yes</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p><bold>Note:</bold> Estimated Training Compute:Stable Diffusion: 1.5&#x007E;150,000 A100 h (based on v1-x estimation); Stable Diffusion XL: SD 2.0 estimated at 0.2M A100 h</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Resource efficiency and hardware demands of state-of-the-art image generative models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th align="left">Std. VRAM Req.</th>
<th align="left">Opt. Min. VRAM</th>
<th align="left">Optimization Techniques</th>
</tr>
</thead>
<tbody>
<tr>
<td>Stable diffusion 1.5</td>
<td>4&#x2013;6 GB</td>
<td>&#x007E;2 GB</td>
<td>Attention slicing, low-VRAM Mode, Core ML</td>
</tr>
<tr>
<td>Stable diffusion XL</td>
<td>8 GB</td>
<td>&#x007E;6 GB</td>
<td>FP8 Quantization, CPU offloading, Microsoft olive</td>
</tr>
<tr>
<td>Stable diffusion 3.0</td>
<td>12&#x2013;24 GB (est.)</td>
<td>Not specified</td>
<td>FP16 quantization, Model offloading (assumed)</td>
</tr>
<tr>
<td>Stable diffusion 3.5</td>
<td>12&#x2013;24 GB</td>
<td>&#x007E;8 GB</td>
<td>FP8 quantization, CPU offloading, 8-bit DiT</td>
</tr>
<tr>
<td>FLUX.1</td>
<td>24 GB&#x002B;</td>
<td>6 GB</td>
<td>FP8/NF4 quantization, System RAM offloading</td>
</tr>
<tr>
<td>HunyuanDiT V1.1</td>
<td>12&#x2013;32 GB</td>
<td>&#x007E;8 GB</td>
<td>Q4/Q5 GGUF, TensorRT, Flash-attention</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s3_1">
<label>3.1</label>
<title>Stable Diffusion Models</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Image Synthesis via Stable Diffusion Models</title>
<p>Stable Diffusion, a prominent tool in anime image synthesis, leverages latent diffusion to generate high-fidelity images through iterative denoising in latent space. Developed by StabilityAI, CompVis, and Runway, its initial release in October 2022 marked a significant advancement. Subsequent iterations, including Stable Diffusion XL 1.0 (July 2023), Stable Diffusion 3.0 (February 2024), and Stable Diffusion 3.5 (October 2024), have progressively enhanced resolution, text alignment, and overall generative performance. Notably, the FLUX model (August 2024) [<xref ref-type="bibr" rid="ref-112">112</xref>], by Black Forest Labs, presents a competitive alternative, demonstrating superior image quality. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> illustrates the evolution of image generation quality across Stable Diffusion versions.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Evolution of Stable Diffusion Models demonstrated via images generated under consistent input conditions. See <xref ref-type="app" rid="app-1">Appendix A</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-6.tif"/>
</fig>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Ancillary Technologies in Stable Diffusion Ecosystems</title>
<p>The open-source nature of Stable Diffusion has fostered a diverse ecosystem of ancillary technologies, enhancing its applicability in artistic creation.</p>
<p><bold>Controllable Image Synthesis:</bold> ControlNet (February 2023) enables precise pose and detail manipulation through trainable copies conditioned on edge and contour maps [<xref ref-type="bibr" rid="ref-78">78</xref>], As demonstrated in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. T2I-adapter (February 2023) provides analogous functionality with a lightweight architecture [<xref ref-type="bibr" rid="ref-113">113</xref>]. MaskDiffusion (March 2024) refines textual controllability [<xref ref-type="bibr" rid="ref-114">114</xref>].</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>ControlNet controls the generation of images by adding constraints</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-7.tif"/>
</fig>
<p><bold>Model Fine-Tuning and Optimization:</bold> LoRA (June 2021) achieves efficient fine-tuning by reducing trainable parameters [<xref ref-type="bibr" rid="ref-73">73</xref>]. Latent Consistency Models (LCMs) (October 2023) accelerate fine-tuning by learning latent space mappings [<xref ref-type="bibr" rid="ref-115">115</xref>].</p>
<p><bold>User Interface and Workflow Management:</bold> stable-diffusion-webui provides a user-friendly web interface for text-to-image synthesis [<xref ref-type="bibr" rid="ref-116">116</xref>]. ComfyUI offers a node-based interface for advanced customization and pipeline construction [<xref ref-type="bibr" rid="ref-117">117</xref>].</p>
<p><bold>Advanced Image Manipulation:</bold> PaintsUndo (August 2024) simulates painting brushstrokes and enables sketch extraction and anime-style transformations [<xref ref-type="bibr" rid="ref-118">118</xref>]. Its generation process is detailed in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>PaintsUndo model generation process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Other Models and Related Technological Landscape</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Notable High-Impact Models</title>
<p>While these models collectively push the boundaries of generative capabilities, the landscape of text-to-image synthesis, beyond established open-source models like Stable Diffusion, is populated by a diverse array of other high-performing architectures.</p>
<p>ImagenFX (Google Labs), leveraging Imagen 2 and 3 [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-119">119</xref>], provides global users with high-fidelity visual content generation through an intuitive interface and robust semantic interpretation. Its integration with Gemini enhances multilingual support and complex scene rendering. Lumina-Image 2.0 (Alpha-VLLM) employs a 2.6B-parameter DiT architecture and a Gemma-2B text encoder [<xref ref-type="bibr" rid="ref-120">120</xref>], demonstrating superior text-following performance in DPG and GenEval benchmarks [<xref ref-type="bibr" rid="ref-121">121</xref>&#x2013;<xref ref-type="bibr" rid="ref-123">123</xref>], particularly for Sino-Anglic prompts. CogView3 (Zhipu AI) utilizes a cascaded diffusion framework and relay diffusion techniques [<xref ref-type="bibr" rid="ref-124">124</xref>], achieving accelerated inference and enhanced generative quality via a three-stage generation strategy and diffusion distillation optimization. Adobe Firefly [<xref ref-type="bibr" rid="ref-106">106</xref>], integrated within Adobe Creative Cloud and powered by Adobe Sensei, offers comprehensive image, video, and audio synthesis, trained on commercially compliant datasets, excelling in 1080p video generation. DALL&#x00B7;E 3 [<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-125">125</xref>], through its integration with GPT, enhances intelligent image generation and editing, improving user interaction naturalness and editing precision. Ideogram focuses on generating images with legible textual elements, addressing a critical limitation in existing generative AI tools [<xref ref-type="bibr" rid="ref-107">107</xref>]. Its rapid synthesis and diverse stylistic options facilitate the creation of high-quality visuals for logo, poster, and graphic design.</p>
<p>Illustrious (OnomaAI Research) [<xref ref-type="bibr" rid="ref-126">126</xref>], built upon the SDXL architecture and leveraging Danbooru tags and multi-level captions, specializes in high-fidelity anime and illustration synthesis, excelling in resolution, color gamut, and anatomical accuracy. Pony Diffusion [<xref ref-type="bibr" rid="ref-127">127</xref>], an SDXL-derived anime model, has garnered significant acclaim on Civitai, recognized for its exceptional stylistic adaptation and high-resolution synthesis.</p>
<p>HunyuanDiT (Tencent) employs a DiT architecture with multi-resolution training and a dual-encoder system [<xref ref-type="bibr" rid="ref-52">52</xref>,<xref ref-type="bibr" rid="ref-111">111</xref>], demonstrating superior text-image consistency, subject clarity, and aesthetic fidelity, particularly for Chinese prompts.</p>
<p>These architectures, through continuous innovation, collectively propel the evolution of text-to-image synthesis. Regarding their specific relevance to anime generation, many proprietary models primarily leverage their robust general prior knowledge to infer and render anime aesthetics. In contrast, models such as Illustrious and Pony Diffusion exemplify targeted specialization, having been comprehensively fine-tuned on anime-specific datasets to significantly enhance the anime generation capabilities of their Stable Diffusion base, achieving remarkable stylistic adaptation and fidelity. It is also noteworthy that HunyuanDiT, an open-source model from Tencent built on the Diffusion Transformer (DiT) architecture, represents a convergence of language model and diffusion model strengths. However, as a comparatively recent entrant to the open-source domain, its subsequent development has often involved the adaptation and integration of techniques originating from the more established Stable Diffusion ecosystem. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> compares images generated by various models under consistent conditions.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Depictions generated by various contemporary large-scale generative models using a consistent textual prompt and parameters. See <xref ref-type="app" rid="app-1">Appendix A</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-9a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-9c.tif"/>
</fig>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Related Technological Developments</title>
<p>Recent advancements significantly enhance <bold>anime generation</bold> by reducing costs, boosting efficiency, and refining control. <bold>&#x201C;Stretching Each Dollar&#x201D;</bold> democratizes high-quality generation through resource-frugal training [<xref ref-type="bibr" rid="ref-128">128</xref>], effective delayed patch masking, and synthetic data incorporation. While democratizing access, it struggles with precise text rendering and granular object control, particularly at high masking rates. <bold>SANA</bold> excels in high-resolution synthesis (up to 4K) via innovations like the deep compression autoencoder (AE-F32) and Linear DiT, improving prompt adherence with complex instructions through a decoder-only LLM [<xref ref-type="bibr" rid="ref-129">129</xref>]. Its speed and consumer hardware deployability lower entry barriers, but challenges remain in guaranteed content safety, controllability, and artifact handling in complex areas like faces and hands, showing a trade-off with raw reconstruction quality. Both accelerate rapid iteration in concept design and storyboarding; &#x201C;Stretching Each Dollar&#x201D; aids smaller studios, while SANA provides high-resolution assets and precise text-to-image alignment for diverse anime aesthetics.</p>
<p><bold>ThinkDiff</bold> and <bold>DREAM ENGINE</bold> advance multimodal information fusion [<xref ref-type="bibr" rid="ref-130">130</xref>,<xref ref-type="bibr" rid="ref-131">131</xref>]. ThinkDiff integrates VLM outputs with diffusion processes for in-context reasoning, enabling complex instruction interpretation and generation based on inferred relationships. Despite being lightweight and robust, it currently lacks high image fidelity and mastery of the full spectrum of complex reasoning tasks. DREAM ENGINE offers efficient text-image interleaved control through a versatile multimodal encoder and a two-stage training regimen, excelling in object-driven generation, complex composition, and free-form image editing. These capabilities are transformative for animation: ThinkDiff could automate storyboarding and character development, while DREAM ENGINE provides precise control for detailed scenes, character refinement, and visual consistency.</p>
<p>For fine-grained control and targeted outputs, <bold>MangaNinja</bold> and <bold>PhotoDoodle</bold> are pivotal [<xref ref-type="bibr" rid="ref-132">132</xref>,<xref ref-type="bibr" rid="ref-133">133</xref>]. MangaNinja provides user-controllable diffusion-based manga line-art colorization, handling discrepancies between references and line art with a dual-branch structure and point-driven control. It enhances color consistency and detail, robust for complex scenarios, yet semantic ambiguity can arise with intricate line art, and a dependency on reference imagery persists. PhotoDoodle introduces an instruction-guided framework for learning artistic image editing from few-shot examples using an EditLoRA module. This enables efficient style capture from minimal data (30-50 examples) for seamless integration and mask-free instruction-based editing, though paired dataset collection and training are practical considerations. These tools significantly streamline labor-intensive processes: MangaNinja improves colorization efficiency and character consistency, while PhotoDoodle offers powerful stylistic application and precise, text-prompted modifications.</p>
<p><bold>FluxSR</bold> optimizes image generation for practical applications like super-resolution via efficient single-step inference through Flow Trajectory Distillation (FTD) [<xref ref-type="bibr" rid="ref-134">134</xref>]. Built on powerful pre-trained text-to-image diffusion, it achieves superior perceptual quality and fidelity in recovering high-frequency details. Despite its high computational cost from a large parameter count and residual periodic artifacts, FluxSR offers transformative potential for enhancing low-resolution anime assets. This investment is justified by its ability to significantly streamline production through efficient upscaling of intermediate frames, ultimately contributing to a more rapid and less labor-intensive workflow.</p>
<p>Finally, <bold>CSD-MT</bold> reduces reliance on large labeled datasets through unsupervised content-style decoupling for facial content and makeup style manipulation [<xref ref-type="bibr" rid="ref-135">135</xref>]. Its efficiency, minimal parameters, and rapid inference offer flexible controls, yet extreme makeup styles can challenge accurate boundary rendering. CSD-MT&#x2019;s generalization to unseen anime makeup styles presents a pertinent avenue for efficiently designing and transferring makeup styles onto anime characters, streamlining visual development and creative exploration.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Sequential Image Synthesis for Manga Generation</title>
<p>Diffusion models [<xref ref-type="bibr" rid="ref-3">3</xref>], renowned for their efficacy in single-image synthesis, are demonstrating progress in generating coherent image sequences, which is a critical aspect of manga production.</p>
<p><bold>Reference-Guided Synthesis:</bold> IP-Adapter (August 2023) guides diffusion processes using reference images, though with reduced textual prompt controllability [<xref ref-type="bibr" rid="ref-136">136</xref>].</p>
<p><bold>Identity Preservation and Control:</bold> InstantID (January 2024) integrates facial and landmark images with textual prompts [<xref ref-type="bibr" rid="ref-137">137</xref>], imposing semantic and spatial constraints for identity consistency. PhotoMaker (December 2023) encodes multiple identity images into a unified embedding [<xref ref-type="bibr" rid="ref-138">138</xref>], preserving identity information while accommodating diverse identity integration. However, both models exhibit limitations in maintaining clothing and scene consistency across sequences.</p>
<p><bold>Thematic Consistency Across Sequences:</bold> StoryDiffusion (May 2024) achieves thematic coherence within image batches by incorporating consistent self-attention into Stable Diffusion [<xref ref-type="bibr" rid="ref-139">139</xref>], facilitating manga-style narrative sequencing. The sequential narrative imagery synthesized by StoryDiffusion is shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. Nevertheless, minor inconsistencies in character details persist in multi-character scenes.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Sequential narrative imagery synthesized via Storydiffusion</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-10.tif"/>
</fig>
<p><bold>Multimodal Integration and Layout Control:</bold> DiffSensei (December 2024) integrates diffusion-based image generation with multimodal large language models (MLLMs) [<xref ref-type="bibr" rid="ref-140">140</xref>], employing masked cross-attention for seamless character feature incorporation and precise layout control. The MLLM-based adapter enables flexible character expression, pose, and action modifications aligned with textual prompts.</p>
<p><bold>High-Fidelity Personalization:</bold> AnyStory (January 2025) utilizes an &#x201C;encoding-routing&#x201D; approach [<xref ref-type="bibr" rid="ref-141">141</xref>], employing ReferenceNet and an instance-aware subject router, to achieve high-fidelity personalization for single and multiple subjects in text-to-image generation. This approach enhances subject detail preservation and textual alignment.</p>
<p>While models like InstantID, PhotoMaker, and StoryDiffusion have shown promising advancements in generating consistent characters across scenes, minor inconsistencies in character details and clothing still persist in multi-character or long-form sequences, highlighting an area for continued research.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Evaluation Standards</title>
<p>Evaluation of generative image synthesis models bifurcates into automated and human-centric paradigms [<xref ref-type="bibr" rid="ref-142">142</xref>]. Automated metrics further diverge into content-invariant and content-variant assessments, tailored to scenarios with and without ground truth, respectively [<xref ref-type="bibr" rid="ref-143">143</xref>]. <xref ref-type="table" rid="table-7">Table 7</xref> delineates the principal metrics and their salient characteristics.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Evaluation metrics for image generation models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Metric category</th>
<th>Metric name</th>
<th>Description</th>
<th>Applicable scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td>Content variation</td>
<td>Inception Score (IS) [<xref ref-type="bibr" rid="ref-144">144</xref>]</td>
<td>Evaluates quality and diversity using a pre-trained classifier (e.g., InceptionNet), measured by entropy. Lower entropy indicates higher quality.</td>
<td>Unconditional generative models, such as GANs</td>
</tr>
<tr>
<td>Content variation</td>
<td>Fr&#x00E9;chet Inception DIstance (FID) [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>Compares the distribution similarity between generated and real images in the feature space, based on the Fr&#x00E9;chet distance.</td>
<td>Conditional and unconditional generative models</td>
</tr>
<tr>
<td>Content invariance</td>
<td>Learned Perceptual Image Patch SimIlarity (LPIPS) [<xref ref-type="bibr" rid="ref-145">145</xref>]</td>
<td>Computes perceptual similarity using VGGNet feature space, measuring structural changes.</td>
<td>Tasks with ground truth</td>
</tr>
<tr>
<td>Content invariance</td>
<td>Structural Similarity Index Measure (SSIM) [<xref ref-type="bibr" rid="ref-143">143</xref>]</td>
<td>Simulates human visual perception, measuring luminance, contrast, and structural similarity. Range from &#x2212;1 to 1.</td>
<td>Tasks with ground truth</td>
</tr>
<tr>
<td>Content invariance</td>
<td>Peak Signal-to-Noise Ratio (PSNR)</td>
<td>Quantifies reconstruction quality through mean squared error (MSE). Higher values indicate better quality.</td>
<td>Tasks with ground truth</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Beyond these foundational metrics, several nuanced evaluations merit consideration: Precision and Recall quantify the fidelity and completeness of generated samples; the F1 Score harmonizes these measures; Kernel Inception Distance (KID) [<xref ref-type="bibr" rid="ref-146">146</xref>] offers a robust alternative to FID using Maximum Mean Discrepancy; and the CLIP Score [<xref ref-type="bibr" rid="ref-14">14</xref>] evaluates semantic alignment between generated images and text prompts using pre-trained CLIP models, critical for text-to-image frameworks.</p>
<p>The discussion of generative models, particularly those fine-tuned or designed for specific aesthetics like anime, remains crucial. While empirical benchmarks are valuable, it&#x2019;s pertinent to acknowledge that standard metrics, such as the FID calculated with a general pre-trained Inception model, may not perfectly capture the perceptual nuances highly valued within specialized artistic styles. This underscores the im-portance of qualitative assessments and human-centric evaluations in understanding a model&#x2019;s true per-formance and aesthetic appeal, especially for complex and subjective domains like anime image synthesis. Further insights can be gained from analyzing representative sets of samples generated from diverse prompts, which can highlight models&#x2019; varying strengths across multiple evaluation axes and demonstrate progress in producing perceptually appealing and contextually relevant imagery.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Summary of This section</title>
<p>The proliferation of generative models, notably Stable Diffusion and Flux, alongside community-driven plugins, has significantly expedited anime creation. Current creative paradigms primarily manifest in: (1) the rapid generation of evocative imagery within non-realistic domains, such as post-apocalyptic punk and steampunk, to aid creative ideation; (2) the meticulous refinement of model parameters to instantiate personalized stylistic generators, followed by iterative enhancement of initial outputs, encompassing aesthetic optimization, granular detail augmentation, structural anomaly rectification, and ambient modulation; and (3) the utilization of synthesized images as texture mapping repositories. However, if the complexities of AI-generated content copyright are effectively navigated, and the synthesis process is judiciously controlled, iteratively refined, and finalized by human artists with keen aesthetic sensibility, the output of AI models can meet commercial production requirements. Furthermore, ancillary technologies facilitating the decomposition of image painting processes offer pedagogical utilities for education and novice practitioners.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Video Generation</title>
<p>The year 2024 marked a watershed moment in video synthesis, propelled by augmented computational resources and the refined capacity of generative models to process intricate spatiotemporal data. While still in a nascent stage compared to static image generation, this sub-field is rapidly transforming at the research and developmental frontiers. The advent of sophisticated open-source models has pushed the boundaries of fidelity, duration, and controllability. This section delves into prominent generative video synthesis models, reflecting the field&#x2019;s dynamic, albeit nascent, trajectory. Model selection emphasizes recent breakthroughs, novel architectural paradigms addressing core challenges like temporal coherence and computational efficiency, and demonstrable potential for application within anime production workflows. Notable exemplars, including Google&#x2019;s Gemini 2.0, OpenAI&#x2019;s DALL-E and Sora, Midjourney, and Meta&#x2019;s Make-A-Video, underscore the burgeoning technological sophistication within this domain [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-147">147</xref>].</p>
<sec id="s4_1">
<label>4.1</label>
<title>Large Models for Video Generation</title>
<p>As <xref ref-type="table" rid="table-8">Tables 8</xref> and <xref ref-type="table" rid="table-9">9</xref> illustrate, state-of-the-art video generation models involve even greater computational exigencies compared to their image counterparts, demanding substantial parameters, training resources, and VRAM. Addressing these considerable resource requirements is paramount for wider adoption, with ongoing research focusing on developing more efficient architectures and sophisticated optimization strategies to mitigate hardware demands and facilitate integration into complex animation workflows.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Leading video generation large models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model name</th>
<th align="left">Params</th>
<th align="left">Dev. team/Co.</th>
<th align="left">Rel. date</th>
<th align="left">Open source</th>
</tr>
</thead>
<tbody>
<tr>
<td>Veo 3 [<xref ref-type="bibr" rid="ref-148">148</xref>]</td>
<td>Undisc.</td>
<td>Google deepmind</td>
<td>May 2025</td>
<td>No</td>
</tr>
<tr>
<td>Movie gen video [<xref ref-type="bibr" rid="ref-149">149</xref>]</td>
<td>30B</td>
<td>Meta AI</td>
<td>October 2024</td>
<td>No</td>
</tr>
<tr>
<td>Sora [<xref ref-type="bibr" rid="ref-147">147</xref>]</td>
<td>Undisc.</td>
<td>OpenAI</td>
<td>December 2024</td>
<td>No</td>
</tr>
<tr>
<td>LTX-Video [<xref ref-type="bibr" rid="ref-150">150</xref>]</td>
<td>2B/13B</td>
<td>Lightricks</td>
<td>May 2025</td>
<td>Yes</td>
</tr>
<tr>
<td>Mochi 1 [<xref ref-type="bibr" rid="ref-151">151</xref>]</td>
<td>10B</td>
<td>Genmo</td>
<td>October 2024</td>
<td>Yes</td>
</tr>
<tr>
<td>CogVideoX1.5 [<xref ref-type="bibr" rid="ref-152">152</xref>]</td>
<td>5B</td>
<td>Tsinghua Univ., Zhipu AI</td>
<td>November 2024</td>
<td>Yes</td>
</tr>
<tr>
<td>CogVideoX [<xref ref-type="bibr" rid="ref-152">152</xref>]</td>
<td>2B/5B</td>
<td>Tsinghua Univ., Zhipu AI</td>
<td>August 2024</td>
<td>Yes</td>
</tr>
<tr>
<td>HunyuanVideo [<xref ref-type="bibr" rid="ref-153">153</xref>]</td>
<td>13B</td>
<td>Tencent hunyuan</td>
<td>July 2024</td>
<td>Yes</td>
</tr>
<tr>
<td>Cosmos-1.0 [<xref ref-type="bibr" rid="ref-154">154</xref>]</td>
<td>7B/14B</td>
<td>NVIDIA</td>
<td>January 2025</td>
<td>Yes</td>
</tr>
<tr>
<td>SkyReels-V1 [<xref ref-type="bibr" rid="ref-155">155</xref>]</td>
<td>1.3B/13B</td>
<td>Kunlun world wide</td>
<td>February 2025</td>
<td>Yes</td>
</tr>
<tr>
<td>Wan2.1 [<xref ref-type="bibr" rid="ref-156">156</xref>]</td>
<td>1.3B/7B/13B</td>
<td>Alibaba</td>
<td>April 2025</td>
<td>Yes</td>
</tr>
<tr>
<td>Gen-4 [<xref ref-type="bibr" rid="ref-157">157</xref>]</td>
<td>Undisc.</td>
<td>RunwayML</td>
<td>April 2025</td>
<td>No</td>
</tr>
<tr>
<td>Pika 2.2 [<xref ref-type="bibr" rid="ref-158">158</xref>]</td>
<td>Undisc.</td>
<td>Glen pika</td>
<td>February 2025</td>
<td>No</td>
</tr>
<tr>
<td>Ray2 [<xref ref-type="bibr" rid="ref-159">159</xref>]</td>
<td>Undisc.</td>
<td>Luma labs</td>
<td>January 2025</td>
<td>No</td>
</tr>
<tr>
<td>Stable video diffusion 1.1 [<xref ref-type="bibr" rid="ref-160">160</xref>]</td>
<td>1.52</td>
<td>Stablility AI</td>
<td>February 2024</td>
<td>Yes</td>
</tr>
<tr>
<td>Kuaishou Kling 2.1 [<xref ref-type="bibr" rid="ref-161">161</xref>]</td>
<td>Undisc.</td>
<td>Kuaishou</td>
<td>May 2025</td>
<td>No</td>
</tr>
<tr>
<td>Hailuo 02 [<xref ref-type="bibr" rid="ref-162">162</xref>]</td>
<td>Undisc.</td>
<td>MiniMax</td>
<td>June 2025</td>
<td>No</td>
</tr>
<tr>
<td>Pyramid flow [<xref ref-type="bibr" rid="ref-163">163</xref>,<xref ref-type="bibr" rid="ref-164">164</xref>]</td>
<td>Undisc.</td>
<td>Kuaishou tech</td>
<td>October 2024</td>
<td>Yes</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-8fn1" fn-type="other">
<p>Note: Estimated Training Compute: Sora: Open-Sora 2.0 estimated &#x007E;$200 k training cost as a reference; Pyramid Flow: 20.7 k A100 GPU hours</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Resource efficiency and mitigation of VRAM demands in leading video generative models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model</th>
<th>Std. VRAM Req.</th>
<th>Opt. Min. VRAM</th>
<th>Optimization strategies</th>
</tr>
</thead>
<tbody>
<tr>
<td>LTX-Video</td>
<td>&#x2265;24 GB</td>
<td>8 GB&#x002B;</td>
<td>FP8 quantization, multi-GPU parallelization</td>
</tr>
<tr>
<td>Mochi 1</td>
<td>&#x2265;24 GB</td>
<td>&#x2265;12 GB</td>
<td>FP8 scaled quantization, multi-attention backend</td>
</tr>
<tr>
<td>CogVideoX (2B)</td>
<td>18 GB</td>
<td>3.6 GB</td>
<td>INT8 Quantization, torch.compile</td>
</tr>
<tr>
<td>CogVideoX1.5 (5B)</td>
<td>26 GB</td>
<td>4.4 GB</td>
<td>INT8 Quantization, torch.compile</td>
</tr>
<tr>
<td>HunyuanVideo</td>
<td>&#x2265;32 GB</td>
<td>8 GB</td>
<td>Temporal tiling, LoRA</td>
</tr>
<tr>
<td>Cosmos-1.0</td>
<td>&#x2265;80 GB</td>
<td>&#x2265;39 GB (offloaded)</td>
<td>Model offloading strategies</td>
</tr>
<tr>
<td>SkyReels-V1</td>
<td>&#x2265;8 GB (1.3B)</td>
<td>&#x2265;8 GB</td>
<td>FP8 quantization, Multi-GPU parallelization</td>
</tr>
<tr>
<td>Wan2.1</td>
<td>&#x2265;8 GB (1.3B)</td>
<td>&#x2265;8 GB</td>
<td>Wan-VAE architecture</td>
</tr>
<tr>
<td>Stable video diffusion</td>
<td>&#x2265;80 GB</td>
<td>Not specified</td>
<td>Quality/Memory/Speed Trade-offs</td>
</tr>
<tr>
<td>Pyramid flow</td>
<td>12 GB</td>
<td>&#x2265;6 GB</td>
<td>CPU Offloading, Quantization</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Capabilities and Challenges of Open-Source Generators</title>
<p>The advent of sophisticated open-source models has rapidly transformed video synthesis, pushing the boundaries of fidelity, duration, and controllability. A critical assessment of these advancements reveals diverse architectural strategies and performance profiles.</p>
<p>The <bold>WAN model</bold> leverages a dual image-video training paradigm to achieve broad synthesis capabilities, balancing efficiency and scale through its 1.3 and 14B variants. Its innovative spatiotemporal mechanisms facilitate the capture of complex dynamics, with the Streamer method enabling extended video generation at enhanced speeds. Nonetheless, challenges persist in preserving fine details during substantial motion and managing the computational cost associated with larger models.</p>
<p><bold>Hunyuan Video</bold>, the largest open-source model evaluated at over 13 billion parameters, presents a comprehensive framework integrating advanced data curation and scalable training [<xref ref-type="bibr" rid="ref-153">153</xref>]. It excels in generating high-quality videos with precise text-video alignment and robust conceptual generalization. Its capabilities extend to coherent action sequences and localized text generation, partly attributed to large language model integration. While the original text doesn&#x2019;t explicitly detail disadvantages, challenges include the significant computational resources required and difficulties in maintaining perfect consistency in intricate scenarios. Artefacts from tiling during VAE inference require further refinement.</p>
<p><bold>Stable Video Diffusion (SVD)</bold> marks a significant advancement in high-resolution latent video diffusion, underpinned by a systematic data management pipeline that elevates performance through pre-training [<xref ref-type="bibr" rid="ref-160">160</xref>]. SVD demonstrates a strong command of motion and 3D understanding, serving as a potent 3D prior for multi-view synthesis. Its modularity supports fine-tuning for downstream tasks like image-to-video generation and camera control. However, its efficacy is primarily limited to short videos, struggling with extended sequences, and inherent diffusion model characteristics result in slower sampling and high memory demands.</p>
<p>The <bold>Pyramidal Flow Matching</bold> approach enhances video generative modeling efficiency through a unified algorithm that reduces training tokens via temporal pyramids and integrates trajectories into a single Diffusion Transformer. This method supports high-quality video generation up to 10 s but can introduce subtle subject inconsistencies in longer videos. Current limitations include a lack of support for keyframe or video interpolation and an opportunity for improved fidelity to intricate prompts.</p>
<p><bold>LTX-Video</bold> focuses on realtime video latent diffusion by optimizing a Transformer-based model with a high-compression Video-VAE [<xref ref-type="bibr" rid="ref-150">150</xref>]. This enables efficient spatiotemporal attention and high-resolution output directly in pixel space, achieving faster-than-realtime generation speeds. Yet, the high compression inherently limits fine detail representation, and performance can be sensitive to prompt clarity. The model currently focuses on short videos, with domain-specific adaptability remaining largely unexplored.</p>
<p><bold>CogVideoX</bold> employs an expert Transformer within its diffusion model to generate continuous videos up to 10 s with strong text alignment and coherent actions [<xref ref-type="bibr" rid="ref-152">152</xref>]. It utilizes a 3D VAE for improved compression and fidelity, addressing the persistent challenge of flickering in generated video sequences. While scalable, aggressive compression can hinder convergence, and high-quality fine-tuning might slightly diminish semantic capabilities. Achieving long-term consistency with dynamic narratives is a persistent challenge. <bold>SkyReels-A1</bold> is specifically designed for expressive portrait animation using a video diffusion Transformer framework [<xref ref-type="bibr" rid="ref-155">155</xref>]. It excels at transferring expressions and movements while preserving identity, producing realistic animations adaptable to various proportions. The model handles subtle expressions effectively. Nonetheless, identity distortion, background instability, and unrealistic facial dynamics remain challenges, particularly with extreme pose variations.</p>
<p>Finally, the <bold>Cosmos World Foundation Model Platform</bold> provides a framework for building world models for physical AI, generating high-quality 3D consistent videos with accurate physical attributes through diffusion and autoregressive methods. Its open-source nature promotes accessibility. However, these models are in an early stage of development, exhibiting limitations as reliable physical simulators regarding object permanence, contact dynamics, and instruction following. Evaluating physical fidelity is also a significant challenge.</p>
<p>In summation, the open-source video generative landscape is marked by diverse, powerful models. Ongoing research is essential to address current limitations in fidelity, efficiency, and controllability, paving the way for broader applications and continued innovation.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Performance Evaluation and Anime Potential</title>
<p>Recent advancements in open-source video generative models have demonstrated notable superiority over established baselines and various contemporary models in rigorous paper evaluations. The WAN model, for instance, has consistently surpassed existing open-source and sophisticated commercial solutions across a spectrum of internal and external benchmarks, exhibiting a decisive performance advantage, notably outperforming models such as Sora, Hunyuan, and various CN-Top variants in weighted scores and human preference studies. Despite these impressive gains in fidelity and text alignment, challenges persist, particularly in achieving perfect consistency in long-form narratives, maintaining fine detail during significant motion, and fully meeting the high standards for stylistic precision and emotional depth inherent in complex anime productions. Similarly, Hunyuan Video has been shown to exceed the performance of prior state-of-the-art models, including Runway Gen-3 and Luma 1.6, alongside several prominent domestic models, particularly excelling in text alignment, motion quality, and visual fidelity assessments [<xref ref-type="bibr" rid="ref-153">153</xref>]. Stable Video Diffusion (SVD) has proven superior to models like GEN-2 and PikaLabs in image-to-video generation quality and significantly outperformed models such as CogVideo, Make-A-Video, and Video LDM in zero-shot text-to-video generation metrics [<xref ref-type="bibr" rid="ref-160">160</xref>]. The PyramidFlow model has distinguished itself by surpassing all evaluated open-source video generation models on comprehensive benchmarks like VBench and EvalCrafter, achieving parity with commercial counterparts such as Kling and Gen-3 Alpha using exclusively public datasets. LTX-Video has demonstrated a considerable lead over models including Open-Sora Plan, CogVideoX (2B), and PyramidFlow in user preference studies for both text-to-video and image-to-video tasks [<xref ref-type="bibr" rid="ref-144">144</xref>]. CogVideoX-5B has shown dominance over a range of models including T2V-Turbo, AnimateDiff, and VideoCrafter-2.0 in automated evaluations and outperformed the closed-source Kling in human assessments. Furthermore, SkyReels-A1 has exhibited superior generative fidelity and motion accuracy compared to diffusion and non-diffusion models like Follow-Your-Emoji and LivePortrait, also achieving higher image quality than most existing methods [<xref ref-type="bibr" rid="ref-155">155</xref>]. Lastly, the Cosmos World Foundation Model Platform&#x2019;s components have showcased remarkable performance improvements, with its Tokenizer outperforming existing tokenizers like CogVideoX-Tokenizer and Omni-Tokenizer in key metrics, and its World Foundation Models demonstrating significant advantages over VideoLDM and CamCo in 3D consistency, view synthesis, camera control, and instruction-based video prediction [<xref ref-type="bibr" rid="ref-152">152</xref>]. These findings collectively underscore the rapid progress and increasing competitive edge of open-source initiatives in the video generation domain.</p>
<p>Capitalizing on their recent performance breakthroughs, these sophisticated video generative models constitute a significant development for anime production, and can augment the creative and technical palette available to animators. These models demonstrate capabilities ranging from generating diverse artistic styles and handling multi-language text integration to exhibiting robust generalization in avatar animation tasks, including anime and CGI characters, with precise control over pose and expression. The capacity for high-resolution, temporally consistent video generation from text or images provides valuable tools for preliminary concept visualization, storyboarding, and generating certain complex scenes or effects. This can augment workflow efficiency and reduce production overheads. Furthermore, specialized models excelling in expressive portrait animation with accurate facial and body motion transfer, adaptable to varied anatomies and scene contexts, directly address the nuanced demands of character animation in anime. While challenges persist in achieving perfect consistency in long-form narratives, maintaining fine detail during significant motion, and fully meeting the high standards for stylistic precision and emotional depth inherent in anime, the open availability and progressive capabilities of these models offer a fertile ground for developing next-generation animation techniques and tools. Their foundational strengths in high-quality video synthesis, 3D consistency, and increasing controllability suggest a potential for streamlining pipelines, fostering creative exploration, and contributing to the visual lexicon of anime.</p>
<p>Comparative video generation results from various large models, obtained on identical hardware using a uniform text prompt, are presented in <xref ref-type="fig" rid="fig-11">Figs. 11</xref>&#x2013;<xref ref-type="fig" rid="fig-18">18</xref>.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>CogVideoX 1.5 (5B Parameters). See <xref ref-type="app" rid="app-1">Appendix A</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-11.tif"/>
</fig><fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Cosmos-1.0 (7B Parameters)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-12.tif"/>
</fig><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>HunyuanVideo</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-13.tif"/>
</fig><fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>LTX-Video (2B Parameters)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-14.tif"/>
</fig><fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Mochi 1 (10B Parameters)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-15.tif"/>
</fig><fig id="fig-16">
<label>Figure 16</label>
<caption>
<title>Pyramid-Flow</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-16.tif"/>
</fig><fig id="fig-17">
<label>Figure 17</label>
<caption>
<title>SkyReels-V1</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-17.tif"/>
</fig><fig id="fig-18">
<label>Figure 18</label>
<caption>
<title>WAN2.1</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-18.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Contemporary Strides in Video Synthesis</title>
<p>Drawing upon recent breakthroughs in generative AI, particularly within diffusion and language models, the landscape of anime video synthesis is rapidly advancing. A critical challenge in this domain lies in effectively generating extended, multi-scene narratives with both temporal and character consistency, while simultaneously addressing the exponential increase in computational demands Contemporary research endeavors are addressing this complex challenge through innovative architectural designs and optimization strategies. Techniques such as global-local diffusion cascades and segmented cross-attention mechanisms are being developed to enhance efficiency in processing long video sequences. Alongside these architectural improvements, optimization strategies like temporal tiling and parameter-efficient fine-tuning (e.g., LoRA) are crucial for managing computational resources, particularly VRAM. Furthermore, significant strides are being made in explicitly improving temporal coherence through methods like consistent self-attention, temporal-aware positional encoding, and latent state variable-based modeling. Concurrently, advancements in character animation are focusing on maintaining consistent appearances across frames and scenes via appearance encoders, multi-scale feature fusion networks, and query injection mechanisms. While challenges persist, the convergence of these architectural, algorithmic, and optimization-focused innovations is contributing to progress towards more efficient and consistent generation of long-form anime video content. A overview of these advancements is presented in <xref ref-type="fig" rid="fig-19">Fig. 19</xref>.</p>
<fig id="fig-19">
<label>Figure 19</label>
<caption>
<title>Overview of advancements in video generation technologies</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-19.tif"/>
</fig>
<p>Key advancements can be categorized into four interconnected areas, with particular emphasis on their implications for anime:</p>
<p>Diffusion Model Optimization and Architectural Innovations: Foundational to high-fidelity video generation, this area focuses on refining diffusion models for enhanced quality, efficiency, and controllability. Core contributions include hybrid pixel-latent diffusion frameworks (e.g., Show-1), high-fidelity image fine-tuning (e.g., VideoCrafter2), and human preference alignment [<xref ref-type="bibr" rid="ref-165">165</xref>&#x2013;<xref ref-type="bibr" rid="ref-167">167</xref>]. For anime, these improvements directly translate to sharper visuals and more aesthetically pleasing outputs. Efficiency gains, such as parallel generation strategies (e.g., PAR), are crucial for scaling up anime production, which often involves numerous frames [<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
<p>Long Video and Multi-Scene Generation: This domain tackles the critical challenge of maintaining temporal coherence and content consistency across extended narratives, a paramount concern for anime series and films. Pivotal methods include consistent self-attention and semantic motion predictors (e.g., StoryDiffusion), temporal-aware positional encoding, and latent state variable-based modeling (e.g., Owl-1) [<xref ref-type="bibr" rid="ref-139">139</xref>,<xref ref-type="bibr" rid="ref-168">168</xref>,<xref ref-type="bibr" rid="ref-169">169</xref>]. These innovations are directly applicable to ensuring narrative flow and visual continuity in multi-episode anime productions. Efficiency techniques like global-local diffusion cascades (e.g., NUWA-XL) and segmented cross-attention mechanisms (e.g., Presto) are vital for generating long anime sequences without prohibitive computational costs [<xref ref-type="bibr" rid="ref-170">170</xref>,<xref ref-type="bibr" rid="ref-171">171</xref>].</p>
<p>Character Animation and Consistency: Maintaining consistent character appearances and actions across frames and scenes is indispensable for compelling anime. Research here centers on achieving temporal coherence and visual realism. Seminal works utilize video diffusion models with appearance encoders (e.g., MagicAnimate), multi-scale feature fusion networks, and query injection for cross-shot consistency [<xref ref-type="bibr" rid="ref-172">172</xref>&#x2013;<xref ref-type="bibr" rid="ref-174">174</xref>]. These advancements are directly responsible for preserving character identity and fluidity of motion throughout anime narratives.</p>
<p>Artistic Style and Specific Scene Generation: This area is particularly relevant to anime, which relies heavily on distinctive artistic styles. It focuses on achieving stylistic diversity and scene-specific content. Key methodologies involve T2V priors and deformation techniques and reference image-driven style adapters (e.g., StyleCrafter) [<xref ref-type="bibr" rid="ref-175">175</xref>,<xref ref-type="bibr" rid="ref-176">176</xref>]. These enable precise control over the aesthetic qualities of generated anime, allowing for the replication of diverse artistic styles and the creation of highly customized visual content.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Diffusion Model Optimization and Architectural Innovations</title>
<p>The adaptation of diffusion models to video generation, while promising, necessitates addressing inherent complexities related to temporal dynamics, computational efficiency, and user controllability. Current research endeavors are strategically focused on enhancing generation fidelity, optimizing computational throughput, and augmenting user-directed manipulation.</p>
<p>Regarding fidelity, seminal works have introduced hybrid pixel-latent diffusion frameworks (Show-1, October 2023) to balance quality and efficiency [<xref ref-type="bibr" rid="ref-166">166</xref>], refined spatial modules via high-fidelity image fine-tuning (VideoCrafter2, January 2024) [<xref ref-type="bibr" rid="ref-167">167</xref>], and leveraged human preference alignment through novel metrics (VideoDPO, December 2024) [<xref ref-type="bibr" rid="ref-165">165</xref>]. Furthermore, the reconciliation of reconstruction and generation objectives via VA-VAE and LightningDiT (Reconstruction vs. Generation, March 2025) <italic>has shown</italic> improvements in image quality and training efficiency [<xref ref-type="bibr" rid="ref-177">177</xref>].</p>
<p>For efficiency optimization, PAR (December 2024) introduces a parallel generation strategy by discerning dependencies among visual tokens [<xref ref-type="bibr" rid="ref-39">39</xref>], with the aim of augmenting the efficiency of image and video generation while maintaining the quality of autoregressive models.</p>
<p>Controllability has been enhanced through frameworks that enable cinema-level control over objects and cameras (CineMaster, February 2025) [<xref ref-type="bibr" rid="ref-178">178</xref>], rectified flow transformers achieving benchmark performance (Goku, February 2025), and transformer-based diffusion models that improve sample quality (STG, November 2024) [<xref ref-type="bibr" rid="ref-179">179</xref>].</p>
<p>These advancements collectively propel diffusion models towards a synergistic enhancement of quality, efficiency, and controllability, expanding technological frontiers through multimodal integration, spatial-temporal decoupling, and efficient architectural paradigms.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Long Video and Multi-Scene Generation</title>
<p>The generation of long-form and multi-scene videos poses significant challenges related to temporal coherence, content consistency, and computational scalability. Innovations in this domain emphasize consistency preservation, efficiency augmentation, and user-directed control.</p>
<p>Consistency is addressed via consistent self-attention and semantic motion predictors (StoryDiffusion, May 2024) [<xref ref-type="bibr" rid="ref-139">139</xref>], temporal-aware positional encoding (Mind the Time, December 2024) [<xref ref-type="bibr" rid="ref-169">169</xref>], latent state variable-based world evolution modeling (Owl-1, December 2024) [<xref ref-type="bibr" rid="ref-168">168</xref>], and thematic element extraction through cross-modal alignment (Phantom, February 2025) [<xref ref-type="bibr" rid="ref-180">180</xref>].</p>
<p>Efficiency gains are realized through global-local diffusion cascades (NUWA-XL, March 2023) and segmented cross-attention mechanisms (Presto, December 2024) [<xref ref-type="bibr" rid="ref-170">170</xref>,<xref ref-type="bibr" rid="ref-171">171</xref>], enabling the generation of extended video segments with reduced computational overhead.</p>
<p>Controllability is enhanced through LLM-driven multi-scene script generation pipelines (Vlogger, March 2024; VideoStudio, January 2024) [<xref ref-type="bibr" rid="ref-181">181</xref>,<xref ref-type="bibr" rid="ref-182">182</xref>], facilitating complex narrative construction and user-directed content creation.</p>
<p>Despite these advancements, challenges remain in maintaining long-term consistency, mitigating computational demands, and ensuring seamless multi-scene transitions.</p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Artistic Style and Specific Scene Generation</title>
<p>This domain focuses on achieving stylistic diversity and scene-specific content generation, catering to artistic creation and personalized customization. Research avenues encompass style transfer and generation, and specific scene synthesis.</p>
<p>Style transfer is facilitated by T2V priors and deformation techniques (Breathing Life Into Sketches, November 2023) and reference image-driven style adapters (StyleCrafter, September 2024) [<xref ref-type="bibr" rid="ref-175">175</xref>,<xref ref-type="bibr" rid="ref-176">176</xref>], decoupling content and style through pre-training and fine-tuning.</p>
<p>Specific scene generation is enabled by controllable diffusion transformers (VFX Creator, February 2025) [<xref ref-type="bibr" rid="ref-183">183</xref>], allowing for user-directed animated visual effects synthesis.</p>
<p>These methodologies expand the creative potential of video generation through innovative architectural designs and multimodal fusion.</p>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Character Animation and Consistency</title>
<p>Character animation research centers on generating temporally coherent and visually realistic animations. Key objectives include consistency enhancement and controllability augmentation.</p>
<p>Consistency is addressed via video diffusion models and appearance encoders (MagicAnimate, November 2023) [<xref ref-type="bibr" rid="ref-173">173</xref>], multi-scale feature fusion networks and frequency domain stabilization (AnimateAnything, November 2024) [<xref ref-type="bibr" rid="ref-174">174</xref>], and query injection for cross-shot consistency (Multi-Shot Character, December 2024) [<xref ref-type="bibr" rid="ref-172">172</xref>].</p>
<p>Controllability is enhanced through diffusion Transformer-based frameworks (OmniHuman-1, February 2025) and zero-shot, diffusion-based pipelines with dynamic adapters (X-Dyna, January 2025), enabling realistic and contextually rich animations.</p>
<p>Future research aims to address consistency in complex motion and multi-character scenarios and optimize computational efficiency.</p>
</sec>
<sec id="s4_2_5">
<label>4.2.5</label>
<title>Image-to-Video Generation</title>
<p>I2V generation focuses on synthesizing dynamic videos from static images, addressing challenges related to motion inference and content consistency. Research is bifurcated into quality improvement and controllability enhancement.</p>
<p>Quality is improved through text-aligned image context projection and noise connection (DynamiCrafter, November 2023) [<xref ref-type="bibr" rid="ref-184">184</xref>], cascaded models with hierarchical encoders (I2VGen-XL, November 2023) [<xref ref-type="bibr" rid="ref-185">185</xref>], and identity reference networks (Hallo3, March 2025) [<xref ref-type="bibr" rid="ref-186">186</xref>].</p>
<p>Controllability is enhanced via motion field predictors and temporal attention (Motion-I2V, January 2024), spatio-temporal attention and noise initialization (ConsistI2V, July 2024), user-driven cinematic shot design (MotionCanvas, February 2025), and layer-specific control mechanisms (LayerAnimate).</p>
<p>These advancements propel I2V technology towards enhanced realism and user-directed control, though challenges remain in complex scene consistency and computational efficiency.</p>
</sec>
<sec id="s4_2_6">
<label>4.2.6</label>
<title>Audio-Driven Generation</title>
<p>Audio-driven generation synchronizes video content with audio, addressing challenges related to lip synchronization, facial expression control, and multimodal integration. Research is segmented into lip synchronization and expression control, and multimodal fusion and real-time generation.</p>
<p>Lip synchronization is improved via global audio perception and motion decoupled control (Sonic, November 2024), facial motion tokenization (VQTalker, December 2024) [<xref ref-type="bibr" rid="ref-187">187</xref>], memory-guided temporal and emotion-aware audio modules (MEMO, December 2024) [<xref ref-type="bibr" rid="ref-188">188</xref>], and audio-conditional latent diffusion models (LatentSync, December 2024) [<xref ref-type="bibr" rid="ref-189">189</xref>].</p>
<p>Multimodal fusion is enhanced through dual-aspect audio driving (INFP, December 2024) [<xref ref-type="bibr" rid="ref-190">190</xref>], improved patch deletion and noise enhancement (Hallo2, October 2024) [<xref ref-type="bibr" rid="ref-191">191</xref>], explicit motion space and streaming inference (Ditto, June 2024), and two-stage audio-driven virtual avatar generation (EMO2, January 2025) [<xref ref-type="bibr" rid="ref-192">192</xref>].</p>
<p>These technologies advance audio-driven generation towards enhanced realism, multimodal integration, and real-time applicability.</p>
</sec>
<sec id="s4_2_7">
<label>4.2.7</label>
<title>3D and Novel View Generation</title>
<p>This domain focuses on generating 3D models and novel views from 2D inputs, addressing challenges related to 3D structure inference and spatial consistency. Research encompasses 3D model generation and novel view synthesis.</p>
<p>3D model generation is facilitated by bi-modal U-Nets and motion consistency loss (IDOL, December 2024) and SMPL model-guided depth [<xref ref-type="bibr" rid="ref-193">193</xref>], normal, and semantic map fusion (Champ, June 2024) [<xref ref-type="bibr" rid="ref-194">194</xref>].</p>
<p>Novel view synthesis is enhanced by optimized SVD denoising (ViewExtrapolator, November 2024) [<xref ref-type="bibr" rid="ref-195">195</xref>].</p>
<p>These advancements drive the integration of 3D and novel view generation in creative applications.</p>
</sec>
<sec id="s4_2_8">
<label>4.2.8</label>
<title>Other Innovations</title>
<p>Innovations such as physical modeling integration (PhysGen, September 2024) and 3D tracking video-driven diffusion (DaS, January 2025) expand the boundaries of video generation by incorporating domain-specific knowledge [<xref ref-type="bibr" rid="ref-196">196</xref>,<xref ref-type="bibr" rid="ref-197">197</xref>].</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Recent Innovations in Video Editing Methodologies</title>
<p><xref ref-type="fig" rid="fig-20">Fig. 20</xref> provides an overview of recent advancements in video editing.</p>
<fig id="fig-20">
<label>Figure 20</label>
<caption>
<title>Overview of advancements in video editing technologies</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-20.tif"/>
</fig>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Super-Resolution and Depth Estimation</title>
<p>Innovations in video editing are significantly propelled by advancements in super-resolution and depth estimation. STAR (January 2025) leverages a local information enhancement module and dynamic frequency loss [<xref ref-type="bibr" rid="ref-198">198</xref>], coupled with a text-to-video model, to achieve spatio-temporal enhancement in real-world video super-resolution, thereby improving detail fidelity and temporal coherence. Video Depth Anything (January 2025) introduces an efficient spatio-temporal processing head and concise temporal consistency loss [<xref ref-type="bibr" rid="ref-199">199</xref>], enabling high-quality, temporally consistent depth estimation for ultra-long videos, achieving state-of-the-art zero-shot performance.</p>
<p>These methodologies advance video processing by enhancing visual fidelity and geometric perception. STAR balances detail and consistency through multimodal guidance and frequency domain constraints, while Video Depth Anything (January 2025) achieves zero-shot depth estimation with architectural efficiency [<xref ref-type="bibr" rid="ref-199">199</xref>]. Future research must address real-time performance and complex scene adaptability, critical for applications in film production and autonomous driving.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Special Effects Addition and Editing</title>
<p>The domain of special effects addition and editing has seen significant innovation. DynVFX (February 2025) employs a zero-shot, training-free framework [<xref ref-type="bibr" rid="ref-200">200</xref>], utilizing pre-trained text-to-video diffusion and vision-language models, to integrate dynamic content that interacts naturally with scenes based on textual instructions. FramePainter (January 2025) enables interactive image editing through intuitive visual interaction operations [<xref ref-type="bibr" rid="ref-201">201</xref>]. DynamicFace (January 2025) utilizes composable 3D facial priors, diffusion models [<xref ref-type="bibr" rid="ref-202">202</xref>], and temporal layers to achieve high-quality, consistent video face swapping, enhancing identity preservation and expression accuracy.</p>
<p>These advancements augment video/image editing capabilities by facilitating dynamic special effects generation, interactive editing, and identity consistency maintenance. DynVFX lowers the barrier to special effects production [<xref ref-type="bibr" rid="ref-200">200</xref>], FramePainter enhances creative flexibility [<xref ref-type="bibr" rid="ref-201">201</xref>], and DynamicFace resolves coherence challenges in face-swapped videos [<xref ref-type="bibr" rid="ref-202">202</xref>]. Future research must address controllability, complex scene adaptability, and ethical considerations, critical for film, advertising, and virtual content production.</p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Motion Control and Transfer</title>
<p>Motion control and transfer techniques are evolving to enable fine-grained dynamic editing. This shift facilitates the development of content generation tools with enhanced user control, physical plausibility, and usability. Research areas include dynamic parameter adjustment, motion-appearance decoupling, zero-shot motion transfer, 3D-aware motion control, and localized motion customization.</p>
<p>Dynamic parameter adjustment, exemplified by CustomTTT (December 2024) [<xref ref-type="bibr" rid="ref-203">203</xref>], utilizes test-time training to optimize appearance and motion LoRA parameters, mitigating artifact issues in multi-concept combinations. Motion-appearance decoupling, as seen in Motion Modes (November 2024) and MoTrans (December 2024) [<xref ref-type="bibr" rid="ref-204">204</xref>,<xref ref-type="bibr" rid="ref-205">205</xref>], independently models motion and appearance features, enabling decoupled control or transfer. Zero-shot motion transfer, demonstrated by MotionShop (December 2024) and DiTFlow (December 2024) [<xref ref-type="bibr" rid="ref-206">206</xref>,<xref ref-type="bibr" rid="ref-207">207</xref>], transfers motion patterns without target domain data. 3D-aware motion control, as in ObjCtrl-2.5D (December 2024) and Latent-Reframe (December 2024) [<xref ref-type="bibr" rid="ref-208">208</xref>,<xref ref-type="bibr" rid="ref-209">209</xref>], utilizes 3D geometric information to guide motion generation. Localized motion customization, as in MotionBooth (October 2024) [<xref ref-type="bibr" rid="ref-210">210</xref>], enables fine-grained editing of specific regions.</p>
<p>These technologies enhance dynamic controllability through parameter optimization, feature decoupling, zero-shot adaptation, and 3D perception. Future research must address complex interaction modeling, long-term temporal consistency, and ethical implications.</p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Frame Interpolation and In-Betweening</title>
<p>Frame interpolation and in-betweening are pivotal techniques for enhancing video smoothness and visual quality. Diffusion models, as in FILM (July 2022) and VIDIM (April 2024) [<xref ref-type="bibr" rid="ref-211">211</xref>,<xref ref-type="bibr" rid="ref-212">212</xref>], and Transformer architectures, as in MaskINT (April 2024) and EITS (March 2024) [<xref ref-type="bibr" rid="ref-213">213</xref>,<xref ref-type="bibr" rid="ref-214">214</xref>], have significantly advanced this domain. ToonCrafter (May 2024) addresses non-linear motion and occlusion in anime videos through cartoon correction learning and a dual-reference 3D decoder [<xref ref-type="bibr" rid="ref-215">215</xref>].</p>
<p>The integration of diffusion models and Transformers has improved interpolation quality, particularly in complex scenes. Technologies like MaskINT indicate a trend towards real-time applications [<xref ref-type="bibr" rid="ref-213">213</xref>], while ToonCrafter exemplifies the deepening of technology in vertical domains [<xref ref-type="bibr" rid="ref-215">215</xref>].</p>
</sec>
<sec id="s4_3_5">
<label>4.3.5</label>
<title>Video Restoration and Enhancement</title>
<p>Video restoration and enhancement are essential for repairing damaged content and improving video quality. Diffusion models, as in DiffuEraser and SVFR (January 2025) [<xref ref-type="bibr" rid="ref-216">216</xref>,<xref ref-type="bibr" rid="ref-217">217</xref>], attention mechanisms, as in MatAnyone (January 2025) and SeedVR (February 2025) [<xref ref-type="bibr" rid="ref-218">218</xref>,<xref ref-type="bibr" rid="ref-219">219</xref>], and dual-stream architectures, as in VideoPainter (March 2025) [<xref ref-type="bibr" rid="ref-220">220</xref>], have advanced this domain.</p>
<p>These methodologies improve restoration quality and temporal consistency. The integration of multi-tasks, such as super-resolution, restoration, and colorization, and the development of optimization strategies, are key trends.</p>
</sec>
<sec id="s4_3_6">
<label>4.3.6</label>
<title>Quantification of Editing Processes</title>
<p>The MIVE (December 2024) framework quantifies editing leakage through the Cross-Instance Accuracy (CIA) score [<xref ref-type="bibr" rid="ref-221">221</xref>], decoupling Diverse Multi-instance Sampling (DMS) and Instance-center Probability Re-allocation (IPR), achieving state-of-the-art performance.</p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Advanced Paradigms in Video Comprehension</title>
<p>The inherent spatiotemporal complexity and high dimensionality of video data pose significant challenges beyond static image analysis, necessitating models capable of discerning both spatial configurations and intricate temporal dynamics across diverse tasks. While early approaches leveraged 3D Convolutional Neural Networks (3D CNNs) for processing short temporal segments [<xref ref-type="bibr" rid="ref-222">222</xref>], these architectures exhibited inherent limitations in capturing long-range dependencies and demonstrated computational inefficiencies [<xref ref-type="bibr" rid="ref-223">223</xref>]. Subsequently, the advent of Transformer architectures offered global contextual modeling capabilities [<xref ref-type="bibr" rid="ref-64">64</xref>]; however, their quadratic computational complexity with respect to sequence length constrained their applicability to high-resolution or extended-duration video sequences [<xref ref-type="bibr" rid="ref-224">224</xref>].</p>
<p>Recent advancements underscore a discernible trend toward unified, multimodal frameworks. Vitron (October 2024) introduced a pixel-level visual Large Language Model (LLM) capable of comprehensive visual intelligence [<xref ref-type="bibr" rid="ref-225">225</xref>], encompassing understanding, generation, segmentation, and editing across both image and video modalities, thereby addressing a broad spectrum of visual tasks from granular to abstract levels. Similarly, Sa2VA (February 2025) synergistically integrated SAM2 and LLaVA to achieve dense multimodal understanding of both static and dynamic visual content [<xref ref-type="bibr" rid="ref-226">226</xref>], facilitating tasks such as referring segmentation and multimodal dialogue. This trajectory signifies a paradigm shift from unimodal, task-specific architectures towards holistic frameworks that leverage shared representations for cross-modal, multi-task learning, thereby mitigating reliance on bespoke designs and enhancing overall performance.</p>
<p>In the pursuit of enhanced computational parsimony, VideoMamba (April 2025) presented a novel spatiotemporal modeling paradigm characterized by linear complexity [<xref ref-type="bibr" rid="ref-227">227</xref>], effectively circumventing the limitations inherent in both 3D CNNs and Transformer architectures, rendering it particularly efficacious for the analysis of protracted video sequences. Concurrently, architectural innovation is exemplified by Divot (December 2024) [<xref ref-type="bibr" rid="ref-228">228</xref>], a diffusion-based video tokenizer engineered to encapsulate both spatial and temporal feature hierarchies, supporting both video understanding and generative applications. These diverse methodologies not only highlight the heterogeneity of contemporary architectural designs but also reflect an evolving research focus extending beyond mere comprehension towards sophisticated multimodal capabilities.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Evaluation Metrics for Generative Video Synthesis</title>
<p>The evaluation of generative AI video synthesis necessitates comprehensive metrics that quantify spatial quality (frame-level), temporal consistency (cross-frame), and semantic relevance (task adherence). The inherent temporal dimension of video synthesis introduces complexities beyond those encountered in image generation, mandating specialized evaluation paradigms. An overview of common evaluation metrics for video generation models is provided in <xref ref-type="table" rid="table-10">Table 10</xref>.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Evaluation metrics for video generation models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Metric category</th>
<th>Metric name</th>
<th>Description</th>
<th>Applicable scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td>Spatial quality</td>
<td>Inception Score (IS) [<xref ref-type="bibr" rid="ref-144">144</xref>]</td>
<td>Evaluates single-frame quality and diversity using a pre-trained classifier, based on entropy calculation.</td>
<td>Unconditional video generation, frame-level evaluation</td>
</tr>
<tr>
<td>Spatial quality</td>
<td>Fr&#x00E9;chet Video Distance (FVD) [<xref ref-type="bibr" rid="ref-229">229</xref>]</td>
<td>Compares the distribution similarity between generated and real videos using video feature extractors (e.g., I3D).</td>
<td>Conditional and unconditional video generation</td>
</tr>
<tr>
<td>Spatial quality</td>
<td>Structural Similarity Index (SSIM) [<xref ref-type="bibr" rid="ref-143">143</xref>]</td>
<td>Measures frame-level structural similarity, simulating human visual perception.</td>
<td>Tasks with ground truth</td>
</tr>
<tr>
<td>Spatial quality</td>
<td>Peak Signal-to-Noise Ratio (PSNR)</td>
<td>Quantifies frame-level reconstruction quality through mean squared error.</td>
<td>Tasks with ground truth</td>
</tr>
<tr>
<td>Spatial quality</td>
<td>Learned Perceptual Image Patch Similarity (LPIPS) [<xref ref-type="bibr" rid="ref-145">145</xref>]</td>
<td>Evaluates frame-level perceptual similarity using deep features, scalable to videos.</td>
<td>Tasks with ground truth</td>
</tr>
<tr>
<td>Temporal consistency</td>
<td>Frame difference</td>
<td>Computes pixel-level differences between adjacent frames, measuring temporal smoothness.</td>
<td>All video generation tasks</td>
</tr>
<tr>
<td>Temporal consistency</td>
<td>Optical flow consistency [<xref ref-type="bibr" rid="ref-230">230</xref>]</td>
<td>Evaluates inter-frame motion consistency using optical flow estimation.</td>
<td>Dynamic scene generation</td>
</tr>
<tr>
<td>Overall quality</td>
<td>Video Quality Metrics (VQM) [<xref ref-type="bibr" rid="ref-231">231</xref>]</td>
<td>Combines spatial and temporal factors to evaluate overall video quality.</td>
<td>Video streaming quality evaluation</td>
</tr>
<tr>
<td>Cross-modal consistency</td>
<td>CLIP score [<xref ref-type="bibr" rid="ref-232">232</xref>]</td>
<td>Evaluates semantic consistency between video content and text prompts using the CLIP model.</td>
<td>Text-to-video generation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Recent architectural advancements, notably the synergistic integration of autoregressive language models and diffusion models through Diffusion Transformer (DiT) architectures, exemplified by pioneering large models such as Sora, HunyuanDiT, and the high-performing WAN2.1, have significantly propelled the capabilities of text-to-video generation towards enhanced coherence and fidelity. These sophisticated models underscore the critical need for robust evaluation metrics capable of assessing their multifaceted performance. Evaluation methodologies are broadly categorized into automated and human-centric metrics. Automated metrics encompass spatial, temporal, and cross-modal assessments.</p>
<p>Other Relevant Metrics:</p>
<p>Mean Opinion Score (MOS) [<xref ref-type="bibr" rid="ref-233">233</xref>]: A human-centric metric, employing a 1-to-5 scale to quantify subjective quality.</p>
<p>The evaluation of generative video models highlights the diverse strengths and limitations of current architectures. Metrics such as SSIM, PSNR, and LPIPS provide insights into spatial fidelity and perceptual similarity at the frame level, which are crucial for capturing intricate visual details. Temporal consistency, assessed via Optical Flow Consistency and Frame Difference, speaks to the models&#x2019; ability to generate smooth and coherent motion sequences. Cross-modal alignment, indicated by the CLIP Score, underscores the semantic relevance of the generated content to input prompts. While no single model definitively outperforms others across all criteria, continuous advancements in these architectures are leading to significant progress in achieving high-quality, temporally consistent, and semantically accurate visual narratives.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Section Synthesis: Advancements in Video Generation</title>
<p>The year 2024 marked a period of significant advancement in video generation, driven by rapid increases in computational capacity, architectural innovations in generative models, and sophisticated modeling of complex spatio-temporal dynamics. This section has systematically delineated the salient technological advancements across ten pivotal research domains, ranging from model optimization and motion control to long-form video synthesis and artistic stylization, illustrating the comprehensive trajectory from theoretical conception to practical implementation.</p>
<p>Architectural and optimization strategies, particularly within diffusion models, have yielded significant enhancements in generation efficiency and fidelity through techniques such as pixel-latent space fusion, spatial-temporal module decoupling, and human preference alignment [<xref ref-type="bibr" rid="ref-234">234</xref>]. In the domain of motion control and customized generation, fine-grained manipulation of object motion and camera trajectories has been achieved through parameter decoupling, mixed score guidance, and zero-shot transfer methodologies [<xref ref-type="bibr" rid="ref-235">235</xref>]. Long-form and multi-scene video generation has emerged as a critical area of focus, with advancements in self-attention consistency, LLM-driven script generation, and latent state variable modeling addressing temporal logic and entity coherence challenges [<xref ref-type="bibr" rid="ref-139">139</xref>]. Furthermore, artistic style transfer via vector animation adaptation and style decoupling [<xref ref-type="bibr" rid="ref-236">236</xref>], image-to-video synthesis leveraging multi-stage motion prediction and noise optimization [<xref ref-type="bibr" rid="ref-237">237</xref>], frame interpolation utilizing bidirectional diffusion and masked Transformer models [<xref ref-type="bibr" rid="ref-213">213</xref>], character animation integrating optical flow guidance and frequency domain stabilization, audio-driven generation incorporating global perception and emotion modeling [<xref ref-type="bibr" rid="ref-238">238</xref>], 3D and novel view synthesis through joint video and depth map generation [<xref ref-type="bibr" rid="ref-237">237</xref>], and the exploration of multi-task unified models and physical simulation have collectively expanded the technological horizon [<xref ref-type="bibr" rid="ref-196">196</xref>].</p>
<p>Despite these advancements, video generation confronts persistent challenges, including the maintenance of global consistency in ultra-long videos, the accurate simulation of complex physical phenomena, the achievement of real-time generation efficiency in multimodal interactions, and the facilitation of personalized control in artistic creation. Future research directions, including model lightweighting, world model construction, and multimodal fusion, are poised to overcome these limitations, driving innovation in human-computer collaborative creation and immersive experiences.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Music Composition</title>
<p>This section delineates the emergent paradigm of AI-driven music composition, scrutinizing contemporary large-scale generative architectures, sophisticated synthesis methodologies, audio-visual transduction, and rigorous evaluation protocols. In contrast to the mature domain of image synthesis, diffusion-based music generation, though nascent, exhibits an accelerated developmental vector, indicative of its profound potential. This section delineates pivotal AI-driven music composition models, acknowledging this domain&#x2019;s emergent landscape and varied accessibility. Model selection prioritizes architectures exhibiting significant advancements in generating high-fidelity music and audio, focusing on capabilities pertinent to anime soundtracks and sound design, while considering both open-source availability and demonstrated influence or reported performance.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Leading Large Models for Music Generation</title>
<p>The landscape of music and audio generation has witnessed a transformative shift with the advent of large-scale generative models. As delineated in <xref ref-type="table" rid="table-11">Table 11</xref>, several influential architectures have recently emerged, representing the current vanguard and significantly advancing the state of the art in computationally creative auditory content. While the domain of diffusion-based music generation is relatively nascent compared to image synthesis, its developmental trajectory is notably accelerated, indicating its potential. This section provides a critical overview of some of the most impactful large models in this rapidly evolving field, leveraging the provided information to evaluate their capabilities, limitations, and potential applications, particularly within contexts like anime production.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Leading large models for music generation</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model name</th>
<th align="left">Dev. team</th>
<th align="left">Rel. date</th>
<th align="left">Params</th>
<th align="left">Open source</th>
</tr>
</thead>
<tbody>
<tr>
<td>Music AI sandbox [<xref ref-type="bibr" rid="ref-239">239</xref>]</td>
<td>Google DeepMind</td>
<td>April 2025</td>
<td>N/A</td>
<td>No</td>
</tr>
<tr>
<td>Stable audio [<xref ref-type="bibr" rid="ref-240">240</xref>]</td>
<td>Stability AI</td>
<td>April 2024</td>
<td>13.2B</td>
<td>Yes</td>
</tr>
<tr>
<td>MuseNet [<xref ref-type="bibr" rid="ref-241">241</xref>]</td>
<td>OpenAI</td>
<td>April 2019</td>
<td>N/A</td>
<td>No</td>
</tr>
<tr>
<td>Suno AI v4.0 [<xref ref-type="bibr" rid="ref-242">242</xref>]</td>
<td>Suno AI</td>
<td>November 2024</td>
<td>N/A</td>
<td>No</td>
</tr>
<tr>
<td>V2A [<xref ref-type="bibr" rid="ref-243">243</xref>]</td>
<td>Google Deepmind</td>
<td>June 2024</td>
<td>N/A</td>
<td>No</td>
</tr>
<tr>
<td>MusicGen [<xref ref-type="bibr" rid="ref-244">244</xref>]</td>
<td>Meta</td>
<td>January 2024</td>
<td>1.5B</td>
<td>Yes</td>
</tr>
<tr>
<td>Music-01</td>
<td>Minimax</td>
<td>August 2024</td>
<td>N/A</td>
<td>No</td>
</tr>
<tr>
<td>ACE-Step [<xref ref-type="bibr" rid="ref-245">245</xref>]</td>
<td>ACE Studio, JieyueXingchen</td>
<td>May 2025</td>
<td>3.5B</td>
<td>Yes</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Among the prominent models, Stable Audio Open, developed by Stability AI, warrants specific scrutiny. Released in April 2024 with 13.2 billion parameters and notably open-source, this model is primarily oriented towards text-to-audio generation, excelling in crafting high-quality stereo sound effects and field recordings at a professional 44.1 kHz sample rate. This capacity holds considerable promise for enriching the sound design and ambient audio layers in animated productions. Its accessibility, being runnable on consumer-grade GPUs, further democratizes its potential use in creative workflows. Furthermore, its support for generating variable audio lengths (up to 47 s) offers flexibility for diverse scene requirements. However, critical limitations exist: the model currently struggles with generating audio containing connected speech or understandable voice/singing, a significant impediment for anime narratives reliant on dialogue and vocal performances. Additionally, its performance in generating high-quality music is noted as limited compared to some state-of-the-art music-specific models, likely influenced by its training predominantly on Creative Commons licensed data. Empirical evaluation indicates its strength in sound generation (e.g., outperforming several AudioLDM2 variants and AudioGen on AudioCaps FAD openl3) but shows it lags behind other Stable Audio iterations in instrumental music generation on datasets like Song Describer, though it performed slightly better than MusicGen on the latter.</p>
<p>Another pivotal contribution comes from Meta&#x2019;s MusicGen, a 1.5 billion parameter model released in January 2024 as an open-source project. MusicGen distinguishes itself as a single language model architecture capable of generating high-quality monophonic and stereophonic music conditioned on text descriptions or melodic inputs. This conditional capability offers enhanced control over the output. The model employs an efficient token interleaving pattern within a single-stage Transformer, eschewing the need for cascaded models. Extensive empirical evaluation, particularly comprehensive human assessments, positions MusicGen favorably against established baselines such as MusicLM, Mousai, and Riffusion, demonstrating superior output quality and adherence to textual prompts. The model&#x2019;s robust controllability via text and especially melody renders it particularly germane to anime production. The ability to generate scores that align with specific moods dictated by text or to conform to provided melodic structures offers granular control crucial for synchronizing music with on-screen action and narrative flow. Nonetheless, challenges persist, including limitations in achieving fine-grained control without heavy reliance on classifier-free guidance and potential biases stemming from the predominance of Western-style music in its training data.</p>
<p>Beyond these models for which detailed evaluations were provided, <xref ref-type="table" rid="table-11">Table 11</xref> lists several other notable large models contributing to the rapidly evolving field of music generation. These include Google DeepMind&#x2019;s Music AI Sandbox (May 2024) and V2A (June 2024), OpenAI&#x2019;s pioneering MuseNet (April 2019), Suno AI&#x2019;s Suno AI v4.0 (November 2024), and Minimax&#x2019;s Music-01 (August 2024). While detailed technical reports for many of these are not publicly available at the time of writing, their inclusion by prominent research entities underscores the escalating interest and investment in large-scale audio generation. Notably, ACE-Step from ACE Studio and JieyueXingchen, listed with 3.5 billion parameters and marked as open source with a prospective May 2025 release, represents a significant future development in this domain. However, it is important to note that, as of the current date, a formal technical report detailing its architecture and performance metrics is presently unavailable.</p>
<p>Collectively, these leading large models signify a paradigm shift in music and audio generation, moving towards more capable, versatile, and controllable systems. Their diverse strengths and ongoing development trajectories highlight the dynamic evolution of this field, with increasing potential for direct application in creative industries such as anime, enabling more sophisticated and tailored auditory experiences, despite current limitations in areas like nuanced vocal synthesis for some models.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Generative Synthesis Paradigms: Diffusion and Audio-Visual Interplay</title>
<sec id="s5_2_1">
<label>5.2.1</label>
<title>Advancements in Music Generation</title>
<p><xref ref-type="fig" rid="fig-21">Fig. 21</xref> provides a taxonomy of recent advancements in diffusion-based music generation. Diffusion models have significantly advanced music synthesis, achieving high-fidelity audio through progressive denoising and improving quality, controllability, diversity, and human-AI co-creation. They address limitations of traditional methods in balancing sonic fidelity and structural complexity. However, challenges remain in achieving human-level emotional depth, long-duration coherence, precise attribute manipulation, and reducing computational cost.</p>
<fig id="fig-21">
<label>Figure 21</label>
<caption>
<title>A taxonomy of recent advancements in diffusion-based music generation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-21.tif"/>
</fig>
<p>Recent innovations address these limitations. Stable Audio Open [<xref ref-type="bibr" rid="ref-240">240</xref>], Noise2Music [<xref ref-type="bibr" rid="ref-246">246</xref>], DiffRhythm [<xref ref-type="bibr" rid="ref-247">247</xref>], and TangoFlux exemplify the enhanced fidelity and efficiency attainable through diffusion and flow-matching techniques [<xref ref-type="bibr" rid="ref-248">248</xref>]. Multimodal conditioning, as demonstrated by Seed-Music [<xref ref-type="bibr" rid="ref-249">249</xref>], Music ControlNet [<xref ref-type="bibr" rid="ref-250">250</xref>], and JASCO [<xref ref-type="bibr" rid="ref-251">251</xref>], facilitates fine-grained control via text, symbols, and other modalities. YuE and Both Ears Wide Open extend generative capabilities to lyrics-driven full-song synthesis and spatial audio [<xref ref-type="bibr" rid="ref-252">252</xref>,<xref ref-type="bibr" rid="ref-253">253</xref>], respectively. MusicMagus and SMITIN introduce sophisticated editing and intervention techniques [<xref ref-type="bibr" rid="ref-254">254</xref>,<xref ref-type="bibr" rid="ref-255">255</xref>], enabling nuanced manipulation of musical attributes.</p>
<p>Current trajectories underscore the pursuit of efficient, high-quality synthesis (TangoFlux, DiffRhythm) [<xref ref-type="bibr" rid="ref-247">247</xref>,<xref ref-type="bibr" rid="ref-248">248</xref>], personalized control (Music ControlNet, JASCO) [<xref ref-type="bibr" rid="ref-250">250</xref>,<xref ref-type="bibr" rid="ref-251">251</xref>], and multimodal integration (YuE, Seed-Music) [<xref ref-type="bibr" rid="ref-249">249</xref>,<xref ref-type="bibr" rid="ref-252">252</xref>].</p>
</sec>
<sec id="s5_2_2">
<label>5.2.2</label>
<title>Cross-Modal Audio-Visual Synthesis</title>
<p><xref ref-type="fig" rid="fig-22">Fig. 22</xref> provides a taxonomy of recent advancements in cross-modal audio-visual music generation. Traditional music video (MV) production, characterized by intensive interdisciplinary collaboration, suffers from inherent inefficiencies, elevated costs, and challenges in cross-modal alignment. The escalating demand for personalized audio-visual content necessitates the development of efficient, automated, and artistically expressive cross-modal synthesis techniques.</p>
<fig id="fig-22">
<label>Figure 22</label>
<caption>
<title>A taxonomy of recent advancements in cross-modal audio-visual music generation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-22.tif"/>
</fig>
<p>Moving beyond unimodal generative paradigms, AV-Link introduces a unified framework leveraging temporally aligned diffusion features for bidirectional audio-video information exchange [<xref ref-type="bibr" rid="ref-256">256</xref>]. This architecture emulates human audio-visual cognition, facilitating dynamic modulation of inter-modal relationships.</p>
<p>Current research increasingly emphasizes multimodal collaborative control. MMAudio [<xref ref-type="bibr" rid="ref-257">257</xref>], MultiFoley [<xref ref-type="bibr" rid="ref-258">258</xref>], and Wang et al.&#x2019;s work exemplify this trend [<xref ref-type="bibr" rid="ref-259">259</xref>], employing joint training and multimodal conditioning via video, text, and audio to enhance generative fidelity and cross-modal synchronization. This fusion amplifies creative latitude, enabling nuanced output manipulation through composite directives.</p>
<p>Global feature matching, however, often results in semantic coherence devoid of rhythmic concordance. Hierarchical modeling, as demonstrated by Stable-V2A and VidMusician [<xref ref-type="bibr" rid="ref-260">260</xref>,<xref ref-type="bibr" rid="ref-261">261</xref>], addresses this by achieving precise semantic and temporal alignment. DREAM-Talk extends this precision to emotional talking face generation [<xref ref-type="bibr" rid="ref-262">262</xref>], utilizing diffusion models and video-to-video rendering for realistic expression and lip synchronization.</p>
<p>For extended video synthesis, LoVA employs a diffusion Transformer architecture to mitigate temporal inconsistencies in long-duration audio generation [<xref ref-type="bibr" rid="ref-263">263</xref>].</p>
</sec>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Datasets and Evaluation Metrics</title>
<p>The advancement of music generation and audio processing necessitates the development of robust datasets and evaluation paradigms. The FakeMusicCaps dataset addresses the critical issue of synthetic music attribution [<xref ref-type="bibr" rid="ref-264">264</xref>], facilitating audio forensics and mitigating copyright infringement. COCOLA introduces a consistency-oriented contrastive learning framework for music audio representation [<xref ref-type="bibr" rid="ref-265">265</xref>], enabling objective evaluation of harmonic and rhythmic coherence in generative models. The Sound Scene Synthesis Challenge, integrating objective and perceptual metrics, elucidates performance disparities across diverse sound categories and architectures. MixEval-X mitigates inconsistencies and biases in current evaluation methodologies by employing a multimodal benchmark and a mixed-adapt-correct pipeline [<xref ref-type="bibr" rid="ref-266">266</xref>], achieving modality-agnostic evaluation and enhancing alignment with real-world distributions.</p>
<p>Evaluation paradigms for generative music synthesis encompass spatial fidelity (pitch, rhythm), temporal coherence (rhythmic fluency), and semantic relevance (style, emotion) [<xref ref-type="bibr" rid="ref-267">267</xref>]. The inherent temporal complexity and artistic dimensionality of music necessitate specialized evaluation protocols, distinct from those employed in image or video synthesis [<xref ref-type="bibr" rid="ref-268">268</xref>]. A summary of these evaluation metrics is provided in <xref ref-type="table" rid="table-12">Table 12</xref>.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Evaluation metrics for music generation models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Metric category</th>
<th>Metric name</th>
<th>Description</th>
<th>Applicable scenarios</th>
</tr>
</thead>
<tbody>
<tr>
<td>Subjective Eval.</td>
<td>Music auditory test</td>
<td>Evaluation through listeners&#x2019; ability to distinguish AI-generated music from human-composed music, or through ratings based on criteria such as creativity and naturalness.</td>
<td>All music generation tasks</td>
</tr>
<tr>
<td>Subjective Eval.</td>
<td>Visual analysis</td>
<td>Expert analysis of waveform plots, spectrograms (e.g., Rainbowgram), musical scores, etc., to assess musicality.</td>
<td>Scenarios requiring expert analysis</td>
</tr>
<tr>
<td>Objective Eval.</td>
<td>Model-based metrics</td>
<td>Evaluation using metrics such as training loss, accuracy, F1 score, and chord prediction accuracy.</td>
<td>General generative model evaluation</td>
</tr>
<tr>
<td>Objective Eval.</td>
<td>Music Domain Metrics (MDM)</td>
<td>- Pitch and Rhythm: Scale consistency, pitch range, pitch count, etc.&#x003C;br&#x003E;- Harmony: Harmonic consistency, chord entropy, etc.&#x003C;br&#x003E;- Style: Style fitting, content preservation, etc.</td>
<td>Symbolic music (e.g., MIDI) evaluation</td>
</tr>
<tr>
<td>Objective Eval.</td>
<td>Fr&#x00E9;chet Audio Distance (FAD)</td>
<td>Comparison of the similarity between the feature distributions of generated audio and real audio, based on the Fr&#x00E9;chet distance.</td>
<td>Audio music generation evaluation</td>
</tr>
<tr>
<td>Comprehensive Eval.</td>
<td>Subjective &#x002B; objective combination</td>
<td>Integration of listener ratings (e.g., harmony, rhythm) with objective metrics (e.g., polyphony, scale consistency).</td>
<td>Comprehensive evaluation</td>
</tr>
<tr>
<td>Comprehensive Eval.</td>
<td>Heuristic evaluation framework</td>
<td>Quantification of musicality using music theory tools (e.g., Circle of Fifths) and comparison with subjective testing results.</td>
<td>Symbolic music evaluation</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Evaluation methodologies are broadly classified into automated (objective) and human-centric (subjective) assessments [<xref ref-type="bibr" rid="ref-268">268</xref>]. Automated metrics are further delineated into model-based, music domain-specific (MDM), and audio-specific metrics.</p>
<p>Beyond these core metrics, emerging and context-specific paradigms deserve consideration:</p>
<p>Structural Complexity Metrics: These automate the analysis of long-form musical structure by quantifying hierarchical segmentation, repetition rates, and variational patterns [<xref ref-type="bibr" rid="ref-269">269</xref>].</p>
<p><bold>Semantic Content Metrics:</bold> Employing emotion/theme classifiers, these metrics assess the emotional and thematic fidelity of generated music, evaluating classifier performance through data augmentation [<xref ref-type="bibr" rid="ref-270">270</xref>,<xref ref-type="bibr" rid="ref-271">271</xref>].</p>
<p><bold>Audio Quality Metrics:</bold> While less prevalent in music synthesis compared to speech, metrics such as spectrogram similarity and signal-to-noise ratio (SNR) provide crucial insights into sonic fidelity [<xref ref-type="bibr" rid="ref-272">272</xref>,<xref ref-type="bibr" rid="ref-273">273</xref>].</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Game Generation</title>
<sec id="s6_1">
<label>6.1</label>
<title>Basic Framework and Visual Effects of Games</title>
<sec id="s6_1_1">
<label>6.1.1</label>
<title>Playable Game Generation</title>
<p>The advent of Artificial Intelligence Generated Content (AIGC) has catalyzed a paradigm shift in game development, advancing from unimodal image synthesis to complex multimodal video generation. However, the synthesis of interactive, high-fidelity playable games remains a formidable challenge. Current research trajectories encompass: diffusion-based methodologies (GameNGen) [<xref ref-type="bibr" rid="ref-274">274</xref>], integrating generative diffusion models with reinforcement learning for agent training; Transformer-based architectures (Oasis) [<xref ref-type="bibr" rid="ref-275">275</xref>], constructing open-world AI models for holistic game synthesis; DiT-based diffusion models (PlayGen) [<xref ref-type="bibr" rid="ref-276">276</xref>], achieving real-time interaction and accurate simulation of game mechanics; and large-scale foundational world models (Genie 2) [<xref ref-type="bibr" rid="ref-277">277</xref>], enabling the generation of infinitely diverse, controllable 3D environments for embodied agent training. These investigations underscore playable game synthesis as a pivotal AIGC direction, poised for further innovation. This section surveys nascent explorations into AI-driven game generation, concentrating on foundational advancements and integral components within this complex domain. Our selection highlights models and frameworks enabling high-dimensional content synthesis, such as 3D assets and environments, and those demonstrating the synergistic integration of diverse generative AI paradigms to address the inherent complexities of interactive game content creation. These sophisticated generative capabilities, enabling complex, high-dimensional content synthesis and multimodal integration for interactive experiences, demonstrate significant parallels and utility within diverse engineering and scientific fields [<xref ref-type="bibr" rid="ref-1">1</xref>]. <xref ref-type="fig" rid="fig-23">Fig. 23</xref> shows a real-time generated game screen from Oasis.</p>
<fig id="fig-23">
<label>Figure 23</label>
<caption>
<title>Oasis&#x2019;s real-time generated game screen</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_66647-fig-23.tif"/>
</fig>
</sec>
<sec id="s6_1_2">
<label>6.1.2</label>
<title>3D Scene Generation</title>
<p>3D scene synthesis constitutes a critical facet of game generation. Wonderland introduces a novel workflow for efficient high-fidelity 3D scene creation from monocular images [<xref ref-type="bibr" rid="ref-278">278</xref>]. DimensionX extends this [<xref ref-type="bibr" rid="ref-279">279</xref>], generating realistic 3D and 4D scenes via video diffusion. Streetscapes demonstrates the generation of large-scale [<xref ref-type="bibr" rid="ref-280">280</xref>], consistent street-view imagery through autoregressive video diffusion, offering potential applications in urban planning and virtual reality. PaintScene4D generates consistent 4D scenes from textual prompts [<xref ref-type="bibr" rid="ref-281">281</xref>], advancing 4D synthesis. SynCamMaster enhances pre-trained text-to-video models for multi-camera video generation [<xref ref-type="bibr" rid="ref-282">282</xref>], ensuring inter-perspective consistency. Kiss3DGen (March 2025) repurposes 2D diffusion models for efficient 3D asset generation via &#x201C;3D Bundle Images&#x201D; [<xref ref-type="bibr" rid="ref-283">283</xref>]. Light-A-Video (February 2025) mitigates lighting inconsistencies in video relighting through a consistent lighting attention (CLA) module and progressive light fusion (PLF) strategy [<xref ref-type="bibr" rid="ref-284">284</xref>]. AuraFusion360 achieves high-fidelity 360&#x00B0; unbounded scene inpainting via depth-aware masking and adaptive guided depth diffusion [<xref ref-type="bibr" rid="ref-285">285</xref>]. Don&#x2019;t Splat your Gaussians (August 2024) introduces a novel volumetric modeling and rendering approach for scattering and emitting media [<xref ref-type="bibr" rid="ref-286">286</xref>]. Fast3R (January 2025) enhances 3D reconstruction efficiency via a Transformer-based architecture [<xref ref-type="bibr" rid="ref-287">287</xref>]. VideoLifter (January 2025) achieves efficient monocular video-to-3D reconstruction through a segment-based local-to-global strategy [<xref ref-type="bibr" rid="ref-288">288</xref>]. Gaga (December 2024) enables precise open-world 3D scene reconstruction and segmentation using 3D-aware memory banks. Bringing Objects to Life facilitates text-guided 3D object animation [<xref ref-type="bibr" rid="ref-289">289</xref>], preserving object identity. These advancements provide robust technical foundations for game development.</p>
</sec>
<sec id="s6_1_3">
<label>6.1.3</label>
<title>3D Character Synthesis</title>
<p>3D character synthesis represents a significant vector in game generation, focusing on the creation of realistic, animatable 3D character models. AniGS introduces animatable Gaussian avatars from monocular images [<xref ref-type="bibr" rid="ref-290">290</xref>], enabling real-time 3D puppet animation. Consistent Human Image and Video Generation employs spatial conditional diffusion for appearance consistency. Pippo (February 2025) achieves multi-view generation from monocular images via multi-stage training and attention bias [<xref ref-type="bibr" rid="ref-291">291</xref>]. PERSE (December 2024) synthesizes animatable personalized 3D avatars from single portraits [<xref ref-type="bibr" rid="ref-292">292</xref>], supporting facial attribute editing via decoupled latent spaces. These innovations bolster game character generation and animation production.</p>
</sec>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Specialized Game Mechanics via Generative AI</title>
<sec id="s6_2_1">
<label>6.2.1</label>
<title>Autonomous Character Behavior</title>
<p>Autonomous character behavior constitutes a critical research vector in game AI, focusing on the synthesis of realistic and intelligent virtual agents. Current non-player characters (NPCs) often exhibit limitations in seamless game integration. Berkowitz (the director of Curiouser Institute) identified a central challenge: insufficient controllability of large language model (LLM)-driven AI, rendering their behaviors unpredictable and misaligned with game design specifications.</p>
<p>To address these constraints, researchers have explored several avenues. Motion Tracks introduces a 2D trajectory-based action representation and Motion Track Policy [<xref ref-type="bibr" rid="ref-293">293</xref>], enabling imitation learning from human video data for robotic agents. Generative Agents conceptualizes computational agents capable of simulating credible human behaviors [<xref ref-type="bibr" rid="ref-294">294</xref>], exhibiting both individual agency and emergent social interactions within simulated environments. Google DeepMind&#x2019;s SIMA interprets natural language directives and executes tasks within diverse 3D video game contexts [<xref ref-type="bibr" rid="ref-295">295</xref>].</p>
</sec>
<sec id="s6_2_2">
<label>6.2.2</label>
<title>Dynamic Character Customization</title>
<p>Dynamic character customization, specifically outfit alteration, is a salient research area in game AI, enhancing player immersion through personalized avatar appearance. Dynamic Try-On employs dynamic attention mechanisms for virtual video try-on [<xref ref-type="bibr" rid="ref-296">296</xref>], preserving garment detail and temporal consistency during complex motions. This methodology offers a robust framework for realistic character outfit customization in gaming applications.</p>
</sec>
<sec id="s6_2_3">
<label>6.2.3</label>
<title>Immersive Driving Simulation</title>
<p>Immersive driving simulation represents a significant research domain within game AI, aiming to provide authentic and engaging vehicular experiences. The Stag-1 model facilitates the reconstruction of real-world driving scenarios and the synthesis of controllable 4D driving simulations [<xref ref-type="bibr" rid="ref-297">297</xref>], offering a novel approach to autonomous driving simulation. This methodology enhances the realism and interactivity of driving simulations, presenting potential for widespread adoption in gaming environments.</p>
</sec>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Industrial Deployment of Generative AI in Game Development</title>
<p>Despite the rapid proliferation of diffusion models and large language models in research and open-source communities, the comprehensive industrial deployment of generative AI within sectors like gaming and mainstream anime production remains largely nascent. Current implementations are primarily confined to localized artistic workflows or experimental projects. Beyond diffusion-based methodologies, alternative generative techniques are emerging for game asset creation. For instance, Tencent AI Lab&#x2019;s GiiNEX [<xref ref-type="bibr" rid="ref-298">298</xref>], showcased at GDC 2024, exemplifies this trend, achieving a substantial reduction in urban modeling time from five days to 25 min. Microsoft Research&#x2019;s &#x201C;Muse&#x201D; project pioneers the generation of playable gameplay sequences by learning intricate game dynamics and player interactions from extensive gameplay data [<xref ref-type="bibr" rid="ref-299">299</xref>]. Layer AI is refining mobile game design workflows by leveraging durable AI asset generation pipelines that ensure the reliability and consistency of content creation [<xref ref-type="bibr" rid="ref-300">300</xref>]. Scenario AI empowers game developers with an AI-driven platform to generate a plethora of game assets with consistent stylistic control [<xref ref-type="bibr" rid="ref-301">301</xref>], thereby accelerating the art production lifecycle.</p>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Alternative Applications</title>
<sec id="s7_1">
<label>7.1</label>
<title>Generative Narrative Synthesis</title>
<p><xref ref-type="table" rid="table-13">Table 13</xref> presents a list of leading large language models. Despite the rapid evolution and increasing scale of Large Language Models (LLMs), the automated synthesis of coherent and sustained long-form narratives continues to be a significant challenge. While these models exhibit evolving capabilities in generating fluent prose and localized textual segments and demonstrate progress in text coherence, fundamental limitations persist in managing extensive contextual dependencies, maintaining global narrative consistency, executing complex plot trajectories, and authentically representing nuanced affective arcs. These limitations impede the generation of truly cohesive and compelling stories suitable for complex anime narratives. These inherent constraints often lead to semantic drift, where the narrative coherence degrades over extended passages, leading to inconsistencies in plot, character, and theme [<xref ref-type="bibr" rid="ref-323">323</xref>].</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Leading large language models</title>
</caption>
<table>
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Model name</th>
<th align="left">Dev. team</th>
<th align="left">Rel. date</th>
<th align="left">Params</th>
<th align="left">Category</th>
</tr>
</thead>
<tbody>
<tr>
<td>Gemini 2.5 Pro [<xref ref-type="bibr" rid="ref-302">302</xref>]</td>
<td>Google</td>
<td>March 2025</td>
<td>N/A</td>
<td>MultiModal</td>
</tr>
<tr>
<td>Claude 3.7 [<xref ref-type="bibr" rid="ref-303">303</xref>]</td>
<td>Anthropic</td>
<td>February 2025</td>
<td>N/A</td>
<td>Multimodal</td>
</tr>
<tr>
<td>DeepSeek-R1 [<xref ref-type="bibr" rid="ref-304">304</xref>]</td>
<td>SenseTime Res.</td>
<td>January 2025</td>
<td>N/A</td>
<td>Multimodal</td>
</tr>
<tr>
<td>Mistral Large [<xref ref-type="bibr" rid="ref-305">305</xref>]</td>
<td>Mistral AI</td>
<td>July 2024</td>
<td>N/A</td>
<td>Multimodal</td>
</tr>
<tr>
<td>GPT-4.5 [<xref ref-type="bibr" rid="ref-306">306</xref>]</td>
<td>OpenAI</td>
<td>February 2025</td>
<td>N/A</td>
<td>Multimodal</td>
</tr>
<tr>
<td>Gork v4.0 [<xref ref-type="bibr" rid="ref-307">307</xref>]</td>
<td>xAI</td>
<td>July 2025</td>
<td>N/A</td>
<td>Multimodal</td>
</tr>
<tr>
<td>Pixtral Large [<xref ref-type="bibr" rid="ref-308">308</xref>]</td>
<td>MistralAI</td>
<td>November 2024</td>
<td>12B</td>
<td>Multimodal</td>
</tr>
<tr>
<td>Llama 3.2 [<xref ref-type="bibr" rid="ref-309">309</xref>]</td>
<td>Meta</td>
<td>July 2024</td>
<td>1B/3B/11B/90B</td>
<td>LLM</td>
</tr>
<tr>
<td>Mixtral-of-Expert [<xref ref-type="bibr" rid="ref-310">310</xref>]</td>
<td>MistralAI</td>
<td>April 2024</td>
<td>22B</td>
<td>LLM</td>
</tr>
<tr>
<td>Chatglm 4.0plus [<xref ref-type="bibr" rid="ref-311">311</xref>]</td>
<td>Zhipu AI</td>
<td>August 2024</td>
<td>N/A</td>
<td>LLM</td>
</tr>
<tr>
<td>Gemma 3.0 [<xref ref-type="bibr" rid="ref-312">312</xref>]</td>
<td>GoogleDeepmind</td>
<td>March 2025</td>
<td>1B/4B/12B/27B</td>
<td>LLM</td>
</tr>
<tr>
<td>Phi v4.0 [<xref ref-type="bibr" rid="ref-313">313</xref>]</td>
<td>Microsoft</td>
<td>December 2024</td>
<td>14B</td>
<td>LLM</td>
</tr>
<tr>
<td>DBRX [<xref ref-type="bibr" rid="ref-314">314</xref>]</td>
<td>Mosaic AI</td>
<td>March 2024</td>
<td>132B Tot, 36B Act</td>
<td>LLM</td>
</tr>
<tr>
<td>Nemotron-4 [<xref ref-type="bibr" rid="ref-315">315</xref>]</td>
<td>Nvidia</td>
<td>June 2024</td>
<td>340B</td>
<td>LLM</td>
</tr>
<tr>
<td>Hunyuan [<xref ref-type="bibr" rid="ref-316">316</xref>]</td>
<td>Tencent</td>
<td>February 2025</td>
<td>N/A</td>
<td>LLM</td>
</tr>
<tr>
<td>Hunyuan-T1 [<xref ref-type="bibr" rid="ref-317">317</xref>]</td>
<td>Tencent</td>
<td>November 2024</td>
<td>389B Tot, 52B Act</td>
<td>Multimodal</td>
</tr>
<tr>
<td>Doubao [<xref ref-type="bibr" rid="ref-318">318</xref>]</td>
<td>ByteDance</td>
<td>N/A</td>
<td>N/A</td>
<td>LLM</td>
</tr>
<tr>
<td>Kimi K2 [<xref ref-type="bibr" rid="ref-319">319</xref>]</td>
<td>Moonshot AI</td>
<td>N/A</td>
<td>N/A</td>
<td>LLM</td>
</tr>
<tr>
<td>Tongyi Qianwen [<xref ref-type="bibr" rid="ref-320">320</xref>]</td>
<td>Alibaba cloud</td>
<td>N/A</td>
<td>N/A</td>
<td>LLM</td>
</tr>
<tr>
<td>Qwen2 [<xref ref-type="bibr" rid="ref-321">321</xref>]</td>
<td>Alibaba cloud</td>
<td>Sep 2024</td>
<td>0.5B/1.5B/7B/57B</td>
<td>LLM</td>
</tr>
<tr>
<td>MiniMax [<xref ref-type="bibr" rid="ref-322">322</xref>]</td>
<td>MiniMax</td>
<td>January 2025</td>
<td>456B Tot, 45.9B Act</td>
<td>LLM</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Consequently, while LLMs show potential as tools within the creative writing pipeline, facilitating aspects such as genre-specific content generation or initial drafting, their effective deployment in producing complex narratives currently requires a robust, manually augmented, multi-stage process. This workflow typically includes iterative phases of human-led conceptualization, detailed structural outlining, and meticulous post-generation editing and refinement to compensate for the models&#x2019; deficiencies in maintaining sustained narrative integrity.</p>
<p>Contemporary LLMs exhibit varying degrees of proficiency in tackling the complexities of narrative generation. For instance, Gemini 2.0 demonstrates enhanced logical coherence and temporal continuity in generated text compared to previous iterations [<xref ref-type="bibr" rid="ref-324">324</xref>,<xref ref-type="bibr" rid="ref-325">325</xref>], yet further advancements are <italic>still needed</italic> for reliably producing lengthy, intricate narratives without human intervention. Claude 3.7, while often lauded for its capabilities in character development and adherence to explicit instructions, can yield formulaic content and struggles with the intricate demands of sophisticated plot construction over long durations. DeepSeek-R1 showcases commendable prose quality and creative output [<xref ref-type="bibr" rid="ref-304">304</xref>], but is similarly susceptible to logical inconsistencies and deviations from initial narrative directives. Claude 3.5 Sonnet offers robust instruction following and maintains reasonable logical consistency [<xref ref-type="bibr" rid="ref-326">326</xref>], though perhaps at the expense of pronounced creative flair. Even prevalent models like GPT 4.0 [<xref ref-type="bibr" rid="ref-66">66</xref>], while versatile for a range of textual tasks, are often less suited to the demands of novel-length writing due to a tendency towards formulaic structures and only moderate long-range coherence, making them more suitable for shorter-form or professional content generation.</p>
<p>These observations underscore that while notable progress has occurred, particularly in localized text generation and adherence to immediate prompts, the significant challenge of synthesizing extensive, logically consistent, emotionally resonant, and structurally sound narratives without substantial human oversight largely remains unaddressed. Overcoming the persistent issues of semantic drift and maintaining global coherence are crucial challenges for future research in generative narrative synthesis.</p>
</sec>
<sec id="s7_2">
<label>7.2</label>
<title>Virtual Streamers</title>
<p>The evolution of virtual avatars in interactive media has progressed from software-synthesized vocaloids to human-operated VTubers. A new phase is emerging with autonomous AI-driven virtual streamers like Neuro-sama [<xref ref-type="bibr" rid="ref-327">327</xref>]. These entities utilize advanced AI, including deep learning and natural language processing, for real-time interaction, gameplay, and dynamic responses to viewers. Neuro-sama, the first female AI VTuber, debuted on Twitch in December 2022, showcasing fluent conversational engagement powered by large language models. This paradigm shift indicates a trend towards more personalized and intelligent virtual streaming experiences.</p>
<p>Luna AI exemplifies a state-of-the-art autonomous virtual avatar platform, integrating a suite of high-performance AI models, including GPT, Zhipu AI et al. These models facilitate both local and cloud-based deployments [<xref ref-type="bibr" rid="ref-328">328</xref>].</p>
</sec>
</sec>
<sec id="s8">
<label>8</label>
<title>The Ethical and Progressive Arc of Generative Intelligence</title>
<sec id="s8_1">
<label>8.1</label>
<title>Ethical and Security Implications of Generative AI</title>
<p>Generative Artificial Intelligence (Gen-AI) is significantly influencing user safety, digital content creation, and online platform governance. A salient trend is the accelerating transition of Gen-AI from a theoretical construct to a more pervasive technology in research and specialized applications, presenting both transformative opportunities and critical challenges for its comprehensive industrial adoption [<xref ref-type="bibr" rid="ref-329">329</xref>].</p>
<p>Gen-AI&#x2019;s capacity for nuanced natural language understanding renders it a pivotal tool in user safety. Research indicates its efficacy in detecting sophisticated violations, such as phishing and malware, which are challenging for traditional machine learning paradigms [<xref ref-type="bibr" rid="ref-329">329</xref>]. However, the transformative potential of Gen-AI, exemplified by models like GPT-4o and DALL-E 3, requires rigorous ethical scrutiny. While enabling high-fidelity content creation, Gen-AI introduces concerns regarding bias, authenticity, and potential misuse [<xref ref-type="bibr" rid="ref-330">330</xref>]. Online platforms, grappling with voluminous content, require advanced content moderation mechanisms. Researchers are developing collaborative frameworks leveraging annotation disagreement to address the subjectivity and complexity of toxicity detection [<xref ref-type="bibr" rid="ref-331">331</xref>]. Concurrently, they are exploring the legal, ethical, and practical challenges of AI-driven content moderation, emphasizing algorithmic bias, transparency, and accountability [<xref ref-type="bibr" rid="ref-332">332</xref>]. The increasing verisimilitude of AI-generated content necessitates robust detection methodologies. Interviews with Reddit moderators highlight the threats posed by AI-generated content and the limitations of current detection tools [<xref ref-type="bibr" rid="ref-333">333</xref>]. Research is also investigating techniques and challenges in AI-generated text detection, proposing future research directions [<xref ref-type="bibr" rid="ref-334">334</xref>]. The utilization of synthetic data in model training necessitates vigilance against potential biases. Research underscores the importance of synthetic data while acknowledging its inherent challenges and ethical implications [<xref ref-type="bibr" rid="ref-335">335</xref>]. Growing societal concern regarding the regulation of Gen-AI technologies underscores the need for robust governance frameworks. Research suggests parallels between Gen-AI and social media regulation [<xref ref-type="bibr" rid="ref-336">336</xref>]. The global impact of AI necessitates cross-cultural evaluation. Research highlights geographic biases in current studies and calls for broader global perspectives [<xref ref-type="bibr" rid="ref-337">337</xref>]. Population-specific impacts of AI require targeted research and intervention. Studies address the psychological impacts of algorithmic social media on adolescents, advocating for protective measures [<xref ref-type="bibr" rid="ref-338">338</xref>]. Ethical concerns regarding AI-generated text in scientific research, including transparency, bias, and accountability, are also being addressed [<xref ref-type="bibr" rid="ref-339">339</xref>,<xref ref-type="bibr" rid="ref-340">340</xref>]. The scalability of content moderation and the efficacy of AI-driven solutions are subjects of ongoing debate [<xref ref-type="bibr" rid="ref-341">341</xref>].</p>
</sec>
<sec id="s8_2">
<label>8.2</label>
<title>Specific Ethical Considerations in Anime and Cultural Content Production with Generative AI</title>
<p>Beyond the broad ethical implications of Gen-AI, its application within the specialized domain of anime and cultural content creation introduces a distinct set of considerations that warrant meticulous examination. The integration of generative AI into creative workflows, while augmenting production efficiency and fostering some expanded avenues for artistic expression, concurrently presents complex ethical dilemmas concerning authorship, cultural sensitivity, copyright, and the potential impact on traditional artistic communities [<xref ref-type="bibr" rid="ref-342">342</xref>]. Conversely, Chinese jurisprudence, under the Interim Measures, underscores human authorship, yet recent rulings permit copyright for AI-generated content with substantial human creative input, indicating a nuanced approach to co-creation.</p>
<p>Firstly, authorship becomes increasingly ambiguous when AI systems generate or co-create anime content. The traditional understanding of a singular human creator is challenged, raising questions about attribution, creative intent, and responsibility [<xref ref-type="bibr" rid="ref-343">343</xref>,<xref ref-type="bibr" rid="ref-344">344</xref>]. Current policy frameworks are beginning to address this. For instance, the EU AI Act, with full applicability by mid-2026, mandates transparency for generative AI, requiring disclosure that content is AI-generated and labeling of specific files [<xref ref-type="bibr" rid="ref-345">345</xref>]. Similarly, the UK&#x2019;s consultation proposes transparency from AI firms regarding training data and AI-generated content labeling [<xref ref-type="bibr" rid="ref-346">346</xref>]. However, these frameworks do not fully resolve the philosophical question of creative intent when AI co-creates, leaving a significant gap in defining ultimate accountability. China&#x2019;s Interim Measures (Article 4) mandate adherence to societal morality and socialist values, requiring AI-generated content to avoid cultural insensitivity and objectionable material (Article 10), albeit with broad guidelines for anime.</p>
<p>Secondly, the pervasive nature of anime as a global cultural phenomenon necessitates acute attention to cultural sensitivity. Generative AI models, trained on vast datasets that may reflect inherent biases, risk perpetuating stereotypes, misrepresenting cultural nuances, or inadvertently generating inappropriate content [<xref ref-type="bibr" rid="ref-347">347</xref>,<xref ref-type="bibr" rid="ref-348">348</xref>]. While existing legal frameworks, such as the EU AI Act, do not explicitly address cultural sensitivity, they include transparency requirements for training data summaries, which could indirectly aid in identifying and mitigating biases [<xref ref-type="bibr" rid="ref-345">345</xref>]. The UK&#x2019;s proposals for transparency in training data could also support cultural sensitivity by ensuring AI developers disclose datasets [<xref ref-type="bibr" rid="ref-346">346</xref>]. Conversely, Japan&#x2019;s developer-friendly AI rules, which permit training on copyrighted works without infringement, may exacerbate cultural sensitivity issues if biased data is utilized, as artists face significant challenges in proving misuse [<xref ref-type="bibr" rid="ref-349">349</xref>]. China&#x2019;s Interim Measures (Article 4) mandate adherence to societal morality and socialist values, requiring AI-generated content to avoid cultural insensitivity and objectionable material (Article 10), albeit with broad guidelines for anime.</p>
<p>Thirdly, the intricate landscape of copyright law faces significant challenges with the proliferation of AI-generated anime [<xref ref-type="bibr" rid="ref-350">350</xref>]. Questions arise regarding the copyrightability of AI-generated works, especially if the AI is trained on copyrighted material without explicit permission [<xref ref-type="bibr" rid="ref-351">351</xref>]. Global approaches to this issue diverge considerably. The EU AI Act requires generative AI to comply with EU copyright law and mandates the publication of summaries of copyrighted data used for training [<xref ref-type="bibr" rid="ref-345">345</xref>]. The UK is currently consulting on a copyright exception for AI training on copyrighted materials for commercial purposes, allowing rights holders to reserve rights [<xref ref-type="bibr" rid="ref-346">346</xref>]. In contrast, US copyright law typically requires human authorship, precluding the copyrightability of purely AI-generated works [<xref ref-type="bibr" rid="ref-352">352</xref>]. Japan&#x2019;s 2018 copyright law revision permits AI training on copyrighted works as &#x201C;information analysis&#x201D; without infringement, a stance that largely benefits developers but has sparked controversy among artists and legal experts due to potential implications for creator rights and proof challenges in court [<xref ref-type="bibr" rid="ref-349">349</xref>]. In China, courts recognize copyright for AI-generated works with human creative input, and the Interim Measures (Article 9) compel AI service providers to comply with IP law and ensure lawful training data.</p>
<p>Finally, the burgeoning capabilities of generative AI pose a significant potential impact on traditional artistic communities within the anime industry [<xref ref-type="bibr" rid="ref-353">353</xref>]. While AI can democratize content creation, there is a legitimate concern that it could lead to job displacement for traditional animators, illustrators, and voice actors [<xref ref-type="bibr" rid="ref-354">354</xref>&#x2013;<xref ref-type="bibr" rid="ref-356">356</xref>]. Policies are evolving to mitigate this. The EU AI Act&#x2019;s emphasis on transparency and copyright compliance aims to protect creators by ensuring fair use of their work. Similarly, the UK&#x2019;s proposals underscore creator control and payment, supporting licensing deals to ensure artists can derive revenue from AI utilization [<xref ref-type="bibr" rid="ref-345">345</xref>,<xref ref-type="bibr" rid="ref-346">346</xref>]. However, Japan&#x2019;s more developer-friendly rules may exacerbate job displacement, as evidenced by instances such as Netflix&#x2019;s use of AI-generated background art, which generated global criticism for displacing animators [<xref ref-type="bibr" rid="ref-349">349</xref>,<xref ref-type="bibr" rid="ref-357">357</xref>]. The balance between fostering innovation and safeguarding artist livelihoods remains a contentious issue in policy development, underscoring the necessity for nuanced regulations. While China&#x2019;s Interim Measures (Article 5) foster AI innovation in cultural production, they do not explicitly address job displacement; however, the judicial emphasis on human authorship offers some indirect safeguard for traditional artists, reflecting a balance between innovation and IP protection.</p>
</sec>
<sec id="s8_3">
<label>8.3</label>
<title>Review of Progress and Foundational Challenges</title>
<p>Significant progress has been achieved across various domains of generative artificial intelligence. In image synthesis, diffusion models, exemplified by Stable Diffusion and its successors (e.g., SDXL 1.0, Stable Diffusion 3.0, FLUX) [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>,<xref ref-type="bibr" rid="ref-67">67</xref>], have reached a degree of maturity, demonstrating considerable utility in artistic workflows through enhanced resolution and controllability via mechanisms like ControlNet and LoRA [<xref ref-type="bibr" rid="ref-73">73</xref>,<xref ref-type="bibr" rid="ref-250">250</xref>]. Nevertheless, a persistent technical challenge lies in maintaining rigorous inter-frame consistency in sequential image generation, particularly within complex, multi-character scenes. Standard evaluation paradigms, including Inception Score and Fr&#x00E9;chet Inception Distance [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-144">144</xref>], while informative, often inadequately capture these temporal coherence issues.</p>
<p>The year 2024 marked a significant period for video synthesis, with the emergence of large-scale architectures such as Veo [<xref ref-type="bibr" rid="ref-269">269</xref>], Sora [<xref ref-type="bibr" rid="ref-147">147</xref>], and Movie Gen Video [<xref ref-type="bibr" rid="ref-149">149</xref>]. These models have substantially advanced generative quality and parametric control. Optimization strategies, including pixel-latent space fusion and spatio-temporal module decoupling [<xref ref-type="bibr" rid="ref-166">166</xref>,<xref ref-type="bibr" rid="ref-167">167</xref>], along with architectural innovations aimed at computational efficiency [<xref ref-type="bibr" rid="ref-227">227</xref>,<xref ref-type="bibr" rid="ref-228">228</xref>], have propelled the field forward. Crucially, however, despite advancements in fidelity and motion quality, achieving robust long-duration video consistency and real-time performance remains a significant bottleneck, demanding further exploration of advanced architectural paradigms. These limitations underscore that while the technology shows transformative potential, practical integration into fast-paced production pipelines requires overcoming these specific technical hurdles.</p>
<p>In music composition, diffusion-based models like Stable Audio Open and Noise2Music have elevated audio fidelity and control [<xref ref-type="bibr" rid="ref-240">240</xref>,<xref ref-type="bibr" rid="ref-246">246</xref>]. Yet, considerable hurdles persist in generating extended compositions with profound emotional depth and intricate structural coherence. Furthermore, cross-modal audio-visual transduction, while showing promise in enhancing synchronization, often incurs substantial computational overhead [<xref ref-type="bibr" rid="ref-256">256</xref>,<xref ref-type="bibr" rid="ref-261">261</xref>].</p>
<p>Game generation has witnessed notable strides in synthesizing playable content [<xref ref-type="bibr" rid="ref-274">274</xref>,<xref ref-type="bibr" rid="ref-276">276</xref>], 3D environments, and characters [<xref ref-type="bibr" rid="ref-278">278</xref>,<xref ref-type="bibr" rid="ref-279">279</xref>,<xref ref-type="bibr" rid="ref-290">290</xref>,<xref ref-type="bibr" rid="ref-291">291</xref>]. Despite this, developing interactive AI character behaviors with genuine autonomy and seamless integration remains constrained by limitations in generative fidelity and consistency [<xref ref-type="bibr" rid="ref-293">293</xref>,<xref ref-type="bibr" rid="ref-295">295</xref>].</p>
<p>Narrative synthesis, primarily driven by LLMs [<xref ref-type="bibr" rid="ref-324">324</xref>,<xref ref-type="bibr" rid="ref-326">326</xref>], continues to grapple with generating long-form content that transcends formulaic structures and avoids semantic drift or discontinuous plots. While autonomous virtual avatars highlight the potential for AI-driven real-time interaction [<xref ref-type="bibr" rid="ref-327">327</xref>], the underlying narrative generation capabilities require further maturation.</p>
</sec>
<sec id="s8_4">
<label>8.4</label>
<title>Pragmatic Hurdles in Deploying Generative AI for Anime Production</title>
<p>The assimilation of generative AI into the broader anime production sphere, while theoretically transformative, encounters significant practical challenges that temper its immediate widespread adoption, including considerable economic outlays for high-performance computing infrastructure, the need for specialized technical skills, and the complexities of integrating these tools into existing production pipelines.</p>
<p>Foremost among these are the considerable economic outlays. The initial procurement or development of sophisticated AI models, coupled with the requisite high-performance computing (HPC) infrastructure&#x2014;encompassing powerful GPUs/TPUs and scalable cloud resources&#x2014;represents a significant capital investment. Furthermore, ongoing operational expenditures, including energy consumption, data storage, model maintenance, and software licensing, contribute to a substantial total cost of ownership. However, these significant costs must be weighed against the potential for transformative long-term gains in production velocity, the reduction of highly specialized manual labor, and the unprecedented scalability of content generation that these technologies offer. For larger studios, this represents a strategic investment, while for smaller entities, continued research into more accessible and optimized solutions is crucial to unlock similar benefits [<xref ref-type="bibr" rid="ref-358">358</xref>,<xref ref-type="bibr" rid="ref-359">359</xref>].</p>
<p>Secondly, the infrastructural prerequisites extend beyond mere computational power. Robust data management pipelines are indispensable for curating, processing, and versioning the vast datasets essential for training and fine-tuning bespoke anime-centric AI models. High-bandwidth networking is also critical for efficient data transfer and collaborative workflows. The absence of, or deficiencies in, such infrastructure can significantly limit the utility and scalability of generative AI solutions [<xref ref-type="bibr" rid="ref-360">360</xref>].</p>
<p>Finally, a critical limitation resides in the specialized human capital required. Effective deployment necessitates a multidisciplinary workforce adept in AI/ML, data science, and software engineering, alongside artists and directors skilled in navigating and creatively leveraging these novel tools. A significant skills gap exists between traditional animation competencies and the emerging demands of AI-driven production, necessitating substantial investment in training, upskilling, or the recruitment of specialized talent. The seamless integration of AI into entrenched production pipelines also presents a significant workflow re-engineering challenge, requiring a notable shift in both creative and technical processes [<xref ref-type="bibr" rid="ref-361">361</xref>,<xref ref-type="bibr" rid="ref-362">362</xref>].</p>
<p>Addressing these financial, infrastructural, and skill-based constraints is pivotal for unlocking the full potential of generative AI within the anime industry and transitioning it from a nascent technology to an integral component of mainstream production.</p>
</sec>
<sec id="s8_5">
<label>8.5</label>
<title>Practical Integration and Adoption across Creative Industries</title>
<p>While theoretical advancements in generative AI models for anime production are significant, a critical evaluation of their practical adoption by studios and individual creators reveals a growing trajectory of integration, yielding tangible benefits and novel creative paradigms. The initial computational investments, while substantial, are increasingly offset by enhanced production efficiency, novel artistic expression, and expanded accessibility.</p>
<p>Evidencing this trend, diverse applications span across various creative domains. In visual arts, the award-winning &#x201C;Th&#x00E9;&#x00E2;tre D&#x2019;op&#x00E9;ra Spatial&#x201D; by game designer Jason Allen, refined with Midjourney and Photoshop, underscored AI&#x2019;s capacity as a co-creative tool, even securing a prestigious art competition [<xref ref-type="bibr" rid="ref-363">363</xref>]. The commercial viability and artistic acceptance of AI-generated art are further underscored by significant sales at auctions like Christie&#x2019;s &#x201C;Enhanced Intelligence&#x201D; in March 2025, where pieces such as Jesse Woolston&#x2019;s &#x201C;The Dissolution Waiapu&#x201D; and Holly Herndon &#x0026; Mat Dryhurst&#x2019;s &#x201C;Embedding Study 1 &#x0026; 2&#x201D; commanded substantial prices [<xref ref-type="bibr" rid="ref-364">364</xref>,<xref ref-type="bibr" rid="ref-365">365</xref>]. Quantitative data from platforms like Pixiv (289,064 AI illustrations/comics) and ArtStation (15,557 AI search results) further corroborate the widespread adoption of AI tools among visual artists and illustrators [<xref ref-type="bibr" rid="ref-366">366</xref>,<xref ref-type="bibr" rid="ref-367">367</xref>].</p>
<p>In film and animation, generative AI is transitioning from experimental novelty to integrated workflow. The 2025 film &#x201C;generAIdoscope,&#x201D; entirely conceived and produced using AI for visuals, sound, and score, exemplifies the burgeoning potential for fully AI-driven content creation [<xref ref-type="bibr" rid="ref-368">368</xref>]. More specifically within anime, Frontier Works and KaKa Creation&#x2019;s 2025 experimental series &#x201C;Twins Hinahima&#x201D; showcased how Stable Diffusion could significantly reduce costs in the &#x201C;synthesis and adjustment&#x201D; phases, also addressing common issues like character clipping in Unreal Engine environments [<xref ref-type="bibr" rid="ref-369">369</xref>].</p>
<p>The gaming sector, a close cousin to anime in visual aesthetics, has also seen considerable AI integration. Nvidia&#x2019;s Deep Learning Super Sampling (DLSS) technology leverages AI for intelligent frame generation, enabling higher resolutions and frame rates without prohibitive hardware upgrades, thereby enhancing player experience [<xref ref-type="bibr" rid="ref-370">370</xref>]. Experimental ventures like Anuttacon&#x2019;s &#x201C;Whispers From The Star,&#x201D; spearheaded by miHoYo co-founder Cai Haoyu, validate real-time multimodal AI-driven interaction in games. Furthermore, established IPs are incorporating AI; ZUN, the creator of the &#x201C;Touhou Project,&#x201D; utilized Adobe Firefly for background and texture creation in the 2025 game &#x201C;Touhou Kinjoukyou Fossilized Wonders [<xref ref-type="bibr" rid="ref-371">371</xref>].&#x201D; Even independent developers, as seen with the free 2.5D point-and-click adventure game &#x201C;Echoes of Somewhere,&#x201D; are leveraging AI for generating all in-game art assets, democratizing access to high-quality visuals.</p>
<p>Beyond visual media, AI&#x2019;s influence extends to audio and literature. Pedro Sandoval&#x2019;s 2025 release of two fully AI-generated albums, &#x201C;ZKY-18&#x201D; and &#x201C;Dirty Marilyn,&#x201D; marked a pioneering instance of Spotify-certified AI music. In literature, the 2023 novel &#x201C;Aum Golly 2,&#x201D; co-created by GPT and Midjourney, represents an emergent form of experimental collaborative authorship.</p>
<p>These diverse examples collectively illustrate that generative AI, including diffusion and language models, is not merely a theoretical construct but a practically adopted and increasingly indispensable set of tools across the creative industries. Their integration, while still evolving, is demonstrably enhancing efficiency, fostering novel artistic expressions, and reshaping traditional production pipelines.</p>
</sec>
</sec>
<sec id="s9">
<label>9</label>
<title>Conclusion</title>
<p>This survey has systematically elucidated the transformative impact of diffusion and language models on the landscape of anime generation, charting a trajectory from foundational principles to state-of-the-art applications. Our comprehensive review reveals a field characterized by rapid, albeit uneven, progress. In the domain of image synthesis, models such as Stable Diffusion have achieved considerable maturity, enabling the generation of high-fidelity, stylistically coherent visuals that are already being integrated into artistic workflows. Nevertheless, formidable challenges persist, particularly in maintaining rigorous inter-frame consistency for sequential narratives like manga and mitigating subtle artifacts that detract from professional-grade quality.</p>
<p>The synthesis of higher-dimensional content, particularly video, represents a more nascent yet accelerated frontier. While groundbreaking architectures have significantly enhanced generative fidelity and temporal dynamics, critical bottlenecks remain in achieving robust long-duration spatiotemporal coherence, eliminating visual inconsistencies, and reducing the profound computational exigencies required for training and deployment. Similarly, generative applications in music composition, interactive game creation, and long-form narrative synthesis are still in their emergent phases, grappling with fundamental hurdles related to structural complexity, authentic emotional resonance, semantic integrity, and genuine interactivity.</p>
<p>Looking forward, the continued advancement of this domain is intrinsically predicated upon several key research vectors. The synergistic fusion of autoregressive and diffusion models, particularly through architectures like Diffusion Transformers (DiTs), presents a promising paradigm for enhancing multimodal coherence and controllability. Future work must also prioritize the development of more computationally efficient paradigms, such as model lightweighting and advanced optimization strategies, to democratize access and facilitate integration into practical production pipelines. Furthermore, the construction of sophisticated world models capable of understanding complex causal and physical relationships will be pivotal for achieving the next echelon of realism and interactivity in generated content. The ultimate objective remains the realization of seamless human-computer collaborative creation, where these technologies function as intuitive and powerful extensions of the human artist&#x2019;s creative intent.</p>
<p>Ultimately, realizing this progressive arc of generative intelligence necessitates a concerted research effort that not only pushes architectural and algorithmic boundaries but also rigorously addresses the attendant ethical, legal, and security implications. Issues of copyright, cultural sensitivity, data bias, and the preservation of human creative autonomy are paramount and demand the establishment of robust governance frameworks. Continuous, critical reflection and interdisciplinary collaboration are imperative to navigate these complexities, ensuring that the profound potential of generative AI is harnessed responsibly to enrich, rather than supplant, the vibrant and evolving art form of anime.</p>
</sec>
</body>
<back>
<ack>
<p>The author is grateful for the invaluable support and assistance received from numerous individuals throughout the duration of this project.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was supported by the National Natural Science Foundation of China (Grant No. 62202210).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Yujie Wu, Xing Deng, and Haijian Shao designed the study. Yujie Wu performed the experiments, collected the data, and conducted the initial data analysis. The entire manuscript was drafted by Yujie Wu, with critical revisions and intellectual contributions from Xing Deng and Haijian Shao. Xing Deng and Haijian Shao provided essential guidance and supervision throughout the project, ensuring the integrity and quality of the research. All authors contributed to the final data analysis and interpretation. Ke Cheng and Ming Zhang provided oversight and contributed to the data analysis. Yingtao Jiang and Fei Wang provided expertise in the conceptualization and guidance of the project, and reviewed and edited the final manuscript. Xing Deng is designated as the corresponding author and guarantor of the study. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data sharing is not applicable to this article as no datasets were generated or analyzed during the current study.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref>: Evolution of Stable Diffusion Models demonstrated via images generated under consistent input conditions.</p>
<p>Prompt: masterpiece, best quality, good quality, very awa, newest, highres, absurdres, 1girl, solo, long hair, looking at viewer, blue eyes, shirt, skirt, black hair, long sleeves, closed mouth, jewelry, sitting, white shirt, outdoors, sky, barefoot, cloud, water, from side, bracelet, ocean, scenery, reflection, rock, ruins, bridge, power lines, river, utility pole, evening, reflective water, rubble, limited palette, sketch, glowing, psychedelic, epic composition, epic proportion, dynamic angle, volumetric lighting, masterpiece, best quality, amazing quality, very aesthetic, absurdres, newest, perfect hands, OverCute, FlatLineArt</p>
<p>Checkpoints Models: Stable Diffusion 1.5, Stable Diffusion XL, Stable Diffusion 2.0, Stable Diffusion 3.0, Stable Diffusion 3.5, FLUX 1.0</p>
<p>Hardware: TensorArt Cloud Generation Service</p>
<p><xref ref-type="fig" rid="fig-9">Fig. 9</xref>: Depictions generated by various contemporary large-scale generative models using a consistent textual prompt and parameters.</p>
<p>Prompt: Smooth Quality, 1girl, solo, mature, parted lips, instrument, microphone, meteor shower, starry sky, music, guitar, holding instrument, electric guitar, facing down, messy hair, crazy hair, from below, cowboy shot, epic composition, epic proportion, dynamic angle, masterpiece, best quality, amazing quality, very aesthetic, absurdres, newest, perfect hands, FlatLineArt</p>
<p>Dataset Size: 50 Images</p>
<p>Checkpoints Models: Firefly, CogView3, ImagenFX, Lumina2, HunyuanDit, Ideogram, Illustrious (SDXL), Pony (SDXL), Stable Diffusion 3.5, Niji_style (SD3.5), Flux1, ToxicEchoFlux (FLUX)</p>
<p>Image Generation Hardware: TensorArt Cloud Generation Service, API services provided by relevant commercial companies</p>
<p>Image Testing Hardware: NVIDIA GeForce RTX 4090</p>
<p><xref ref-type="fig" rid="fig-11">Figs. 11</xref>&#x2013;<xref ref-type="fig" rid="fig-18">18</xref></p>
<p>Positive Prompt: Neptune-Hyperdimension Neptunia, 4K FlatLineArt, digital drawing mode, an energetic young girl with light purple hair styled in pigtails and bright purple eyes, wearing her signature purple and white outfit, walking towards the camera from a distance on a brightly lit city street at daytime, full body, dynamic and playful stance, cinematic lighting, perfect anatomy, detailed outfit and D-pad hair ornaments, full HD, 4K, HDR, depth of field</p>
<p>Negative Prompt: bright colors, overexposed, static, blurred details, subtitles, style, artwork, painting, picture, still, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, malformed limbs, fused fingers, still picture, cluttered background, three legs, many people in the background, walking backwards</p>
<p>Checkpoints Models: CogVideoX1.5 5B, Cosmos-1.0-7B, HunyuanVideo, LTX-Video 2B, Mochi 1 10B, Pyramid-Flow-miniFlux, SkyReels-V1, WAN_2_1</p>
<p>Video Generation Hardware: TensorArt Cloud Generation Service, API services provided by relevant commercial companies</p>
<p>Video Testing Hardware: NVIDIA GeForce RTX 4090</p>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Arcucci</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Generative text-guided 3D vision-language pretraining for unified medical image segmentation</article-title>. <comment>arXiv:2306.04811. 2023</comment>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic models</article-title>. <comment>arXiv:2006.11239. 2020</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rombach</surname> <given-names>R</given-names></string-name>, <string-name><surname>Blattmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lorenz</surname> <given-names>D</given-names></string-name>, <string-name><surname>Esser</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ommer</surname> <given-names>B</given-names></string-name></person-group>. <article-title>High-resolution image synthesis with latent diffusion models</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>New Orleans, LA, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>10684</fpage>&#x2013;<lpage>95</lpage>. [cited 2025 Apr 8]. Available from: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content/CVPR2022/html/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper">https://openaccess.thecvf.com/content/CVPR2022/html/Rombach_High-Resolution_Image_Synthesis_With_Latent_Diffusion_Models_CVPR_2022_paper</ext-link>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Pouget-Abadie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Warde-Farley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ozair</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Generative adversarial networks</article-title>. <comment>arXiv:1406.2661. 2014</comment>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nichol</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Classifier guidance: diffusion models beat GANs on image synthesis</article-title>. <comment>arXiv:2105.05233. 2021</comment>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nichol</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shyam</surname> <given-names>P</given-names></string-name>, <string-name><surname>Mishkin</surname> <given-names>P</given-names></string-name>, <string-name><surname>McGrew</surname> <given-names>B</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GLIDE: towards photorealistic image generation and editing with text-guided diffusion models</article-title>. <comment>arXiv:2112.10741. 2022</comment> [cited 2025 Apr 11]. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2112.10741">http://arxiv.org/abs/2112.10741</ext-link>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>MW</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Toutanova</surname> <given-names>K</given-names></string-name></person-group>. <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>. <comment>arXiv:1810.04805. 2019 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1810.04805">http://arxiv.org/abs/1810.04805</ext-link>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Narasimhan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Salimans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Improving language understanding by generative pre-training [Internet]</article-title>. OpenAI; 2018 Jun 11 [cited 2025 Apr 8]. 14 p. Available from: <ext-link ext-link-type="uri" xlink:href="https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf">https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Mok</surname> <given-names>PY</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name></person-group>. <article-title>AnimeDiffusion: anime diffusion colorization</article-title>. <source>IEEE Trans Vis Comput Graph</source>. <year>2024</year>;<volume>30</volume>(<issue>10</issue>):<fpage>6956</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tvcg.2024.3357568</pub-id>; <pub-id pub-id-type="pmid">38261497</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Research on anime-style image generation based on stable diffusion</article-title>. <source>ITM Web Conf</source>. <year>2025</year>;<volume>73</volume>(<issue>4</issue>):<fpage>02038</fpage>. doi:<pub-id pub-id-type="doi">10.1051/itmconf/20257302038</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wagan</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Sidra</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Revolutionizing the digital creative industries: the role of artificial intelligence in integration, development, and innovation</article-title>. <source>SEISENSE J Manag</source>. <year>2024</year>;<volume>7</volume>(<issue>1</issue>):<fpage>135</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.33215/rvcwy166</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Schuhmann</surname> <given-names>C</given-names></string-name>, <string-name><surname>Beaumont</surname> <given-names>R</given-names></string-name>, <string-name><surname>Vencu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gordon</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wightman</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cherti</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LAION-5B: an open large-scale dataset for training next generation image-text models</article-title>. <comment>arXiv:2210.08402. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2210.08402">http://arxiv.org/abs/2210.08402</ext-link>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Heusel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramsauer</surname> <given-names>H</given-names></string-name>, <string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <string-name><surname>Nessler</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hochreiter</surname> <given-names>S</given-names></string-name></person-group>. <article-title>GANs trained by a two time-scale update rule converge to a local nash equilibrium</article-title>. In: <conf-name>Advances in Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017</year> [cited 2025 Apr 8]. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2017/hash/8a1d694707eb0fefe65871369074926d-Abstract.html">https://proceedings.neurips.cc/paper/2017/hash/8a1d694707eb0fefe65871369074926d-Abstract.html</ext-link>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Learning transferable visual models from natural language supervision</article-title>. In: <conf-name>Proceedings of the 38th International Conference on Machine Learning</conf-name>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2021</year> [cited 2025 Apr 8]. p. <fpage>8748</fpage>&#x2013;<lpage>63</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v139/radford21a.html">https://proceedings.mlr.press/v139/radford21a.html</ext-link>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Amankwah-Amoah</surname> <given-names>J</given-names></string-name>, <string-name><surname>Abdalla</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mogaji</surname> <given-names>E</given-names></string-name>, <string-name><surname>Elbanna</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dwivedi</surname> <given-names>YK</given-names></string-name></person-group>. <article-title>The impending disruption of creative industries by generative AI: opportunities, challenges, and research agenda</article-title>. <source>Int J Inf Manag</source>. <year>2024</year>;<volume>79</volume>(<issue>2</issue>):<fpage>102759</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijinfomgt.2024.102759</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname> <given-names>MY</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>AI and animated character design: efficiency, creativity, interactivity</article-title>. <source>Fronti Soc Sci Technol</source>. <year>2024</year>;<volume>6</volume>(<issue>1</issue>). doi:<pub-id pub-id-type="doi">10.25236/fsst.2024.060120</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>The effect of AI on animation production efficiency: an empirical investigation through the network data envelopment analysis</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>24</issue>):<fpage>5001</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13245001</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Salimans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Progressive distillation for fast sampling of diffusion models</article-title>. <comment>arXiv:2202.00512. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2202.00512">http://arxiv.org/abs/2202.00512</ext-link>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nichol</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Improved denoising diffusion probabilistic models</article-title>. <comment>arXiv:2102.09672. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2102.09672">http://arxiv.org/abs/2102.09672</ext-link>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sohl-Dickstein</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Poole</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Score-based generative modeling through stochastic differential equations</article-title>. <comment>arXiv:2011.13456. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2011.13456">http://arxiv.org/abs/2011.13456</ext-link>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Aittala</surname> <given-names>M</given-names></string-name>, <string-name><surname>Aila</surname> <given-names>T</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Elucidating the design space of diffusion-based generative models</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2022</year>;<volume>35</volume>:<fpage>26565</fpage>&#x2013;<lpage>77</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Denoising diffusion implicit models</article-title>. <comment>arXiv:2010.02502. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2010.02502">http://arxiv.org/abs/2010.02502</ext-link>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hoogeboom</surname> <given-names>E</given-names></string-name>, <string-name><surname>Nielsen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jaini</surname> <given-names>P</given-names></string-name>, <string-name><surname>Forr&#x00E9;</surname> <given-names>P</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Argmax flows and multinomial diffusion: learning categorical distributions</article-title>. <comment>arXiv:2102.05379. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2102.05379">http://arxiv.org/abs/2102.05379</ext-link>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Austin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Johnson</surname> <given-names>DD</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tarlow</surname> <given-names>D</given-names></string-name>, <string-name><surname>van den Berg</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Structured denoising diffusion models in discrete state-spaces</article-title>. <comment>arXiv:2107.03006. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2107.03006">http://arxiv.org/abs/2107.03006</ext-link>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Vector quantized diffusion model for text-to-image synthesis</article-title>. <comment>arXiv:2111.14822. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2111.14822">http://arxiv.org/abs/2111.14822</ext-link>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Pseudo numerical methods for diffusion models on manifolds</article-title>. <comment>arXiv:2202.09778. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2202.09778">http://arxiv.org/abs/2202.09778</ext-link>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Analytic-DPM: an analytic estimate of the optimal reverse variance in diffusion probabilistic models</article-title>. <comment>arXiv:2201.06503. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2201.06503">http://arxiv.org/abs/2201.06503</ext-link>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Nichol</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Hierarchical text-conditional image generation with CLIP latents</article-title>. <comment>arXiv:2204.06125. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2204.06125">http://arxiv.org/abs/2204.06125</ext-link>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saharia</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Saxena</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Whang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Denton</surname> <given-names>EL</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Photorealistic text-to-image diffusion models with deep language understanding</article-title>. <source>Adv Neural Inform Process Syst</source>. <year>2022 Dec 6</year>;<volume>35</volume>:<fpage>36479</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pokle</surname> <given-names>A</given-names></string-name>, <string-name><surname>Geng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Kolter</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Deep equilibrium approaches to diffusion models</article-title>. <comment>arXiv:2210.12867. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2210.12867">http://arxiv.org/abs/2210.12867</ext-link>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pfister</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ShadowDiffusion: when degradation prior meets diffusion model for shadow removal</article-title>. <comment>arXiv:2212.04711. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2212.04711">http://arxiv.org/abs/2212.04711</ext-link>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Poole</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Barron</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Mildenhall</surname> <given-names>B</given-names></string-name></person-group>. <article-title>DreamFusion: text-to-3D using 2D diffusion</article-title>. <comment>arXiv:2209.14988. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2209.14988">http://arxiv.org/abs/2209.14988</ext-link>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Takikawa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Magic3D: high-resolution text-to-3D content creation</article-title>. <comment>arXiv:2211.10440. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2211.10440">http://arxiv.org/abs/2211.10440</ext-link>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Latent video diffusion models for high-fidelity long video generation</article-title>. <comment>arXiv:2211.13221. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2211.13221">http://arxiv.org/abs/2211.13221</ext-link>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MedSegDiff: medical image segmentation with diffusion probabilistic model</article-title>. <comment>arXiv:2211.00611. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2211.00611">http://arxiv.org/abs/2211.00611</ext-link>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Singer</surname> <given-names>U</given-names></string-name>, <string-name><surname>Polyak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hayes</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>X</given-names></string-name>, <string-name><surname>An</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Make-a-video: text-to-video generation without text-video data</article-title>. <comment>arXiv:2209.14792. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2209.14792">http://arxiv.org/abs/2209.14792</ext-link>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Salimans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gritsenko</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Norouzi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fleet</surname> <given-names>DJ</given-names></string-name></person-group>. <article-title>Video diffusion models</article-title>. <comment>arXiv:2204.03458. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2204.03458">http://arxiv.org/abs/2204.03458</ext-link>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nichol</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Mishkin</surname> <given-names>P</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Point-E: a system for generating 3D point clouds from complex prompts</article-title>. <comment>arXiv:2212.08751. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2212.08751">http://arxiv.org/abs/2212.08751</ext-link>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Parallelized autoregressive visual generation</article-title>. <comment>arXiv:2412.15119. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.15119">http://arxiv.org/abs/2412.15119</ext-link>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>XL</given-names></string-name>, <string-name><surname>Thickstun</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gulrajani</surname> <given-names>I</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Hashimoto</surname> <given-names>TB</given-names></string-name></person-group>. <article-title>Diffusion-LM improves controllable text generation</article-title>. <comment>arXiv:2205.14217. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2205.14217">http://arxiv.org/abs/2205.14217</ext-link>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Changpinyo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>N</given-names></string-name>, <string-name><surname>Soricut</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Conceptual 12M: pushing web-scale image-text pre-training to recognize long-tail visual concepts</article-title>. <comment>arXiv:2102.08981. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2102.08981">http://arxiv.org/abs/2102.08981</ext-link>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Selvaraju</surname> <given-names>RR</given-names></string-name>, <string-name><surname>Gotmare</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Joty</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hoi</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Align before Fuse: vision and language representation learning with momentum distillation</article-title>. <comment>arXiv:2107.07651. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2107.07651">http://arxiv.org/abs/2107.07651</ext-link>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Rae</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Borgeaud</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>T</given-names></string-name>, <string-name><surname>Millican</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hoffmann</surname> <given-names>J</given-names></string-name>, <string-name><surname>Song</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Scaling language models: methods, analysis &#x0026; insights from training gopher</article-title>. <comment>arXiv:2112.11446. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2112.11446">http://arxiv.org/abs/2112.11446</ext-link>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Podell</surname> <given-names>D</given-names></string-name>, <string-name><surname>English</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lacey</surname> <given-names>K</given-names></string-name>, <string-name><surname>Blattmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dockhorn</surname> <given-names>T</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SDXL: improving latent diffusion models for high-resolution image synthesis</article-title>. <comment>arXiv:2307.01952. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2307.01952">http://arxiv.org/abs/2307.01952</ext-link>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bain</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nagrani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Frozen in time: a joint video and image encoder for end-to-end retrieval</article-title>. <comment>arXiv:2104.00650. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2104.00650">http://arxiv.org/abs/2104.00650</ext-link>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Advancing high-resolution video-language representation with large-scale video transcriptions</article-title>. <comment>arXiv:2111.10337. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2111.10337">http://arxiv.org/abs/2111.10337</ext-link>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>VidProM: a million-scale real prompt-gallery dataset for text-to-video diffusion models</article-title>. <comment>arXiv:2403.06098. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2403.06098">http://arxiv.org/abs/2403.06098</ext-link>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Socher</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>LJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fei-Fei</surname> <given-names>L</given-names></string-name></person-group>. <article-title>ImageNet: a large-scale hierarchical image database</article-title>. In: <conf-name>2009 IEEE Conference on Computer Vision and Pattern Recognition; 2009 Jun 20&#x2013;25</conf-name>; <publisher-loc>Miami, FL, USA</publisher-loc>. p. <fpage>248</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Maire</surname> <given-names>M</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bourdev</surname> <given-names>L</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hays</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Microsoft COCO: common objects in context</article-title>. <comment>arXiv:1405.0312. 2015 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1405.0312">http://arxiv.org/abs/1405.0312</ext-link>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Deep learning face attributes in the wild</article-title>. <comment>arXiv:1411.7766. 2015 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1411.7766">http://arxiv.org/abs/1411.7766</ext-link>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jayasumana</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ramalingam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Veit</surname> <given-names>A</given-names></string-name>, <string-name><surname>Glasner</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chakrabarti</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Rethinking FID: towards a better evaluation metric for image generation</article-title>. <comment>Jun 5]3. 2024 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2401.09603">http://arxiv.org/abs/2401.09603</ext-link>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Peebles</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Scalable diffusion models with transformers</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>. <publisher-loc>Paris, France</publisher-loc>: <publisher-name>IEEE/CVF</publisher-name>; <year>2023 [cited 2025 Apr 8]</year>. p. <fpage>4195</fpage>&#x2013;<lpage>205</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content/ICCV2023/html/Peebles_Scalable_Diffusion_Models_with_Transformers_ICCV_2023_paper.html">https://openaccess.thecvf.com/content/ICCV2023/html/Peebles_Scalable_Diffusion_Models_with_Transformers_ICCV_2023_paper.html</ext-link>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Stability</surname> <given-names>AI</given-names></string-name></person-group>. <article-title>Stable diffusion 3: research paper</article-title>. <comment>[cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://stability.ai/news/stable-diffusion-3-research-paper">https://stability.ai/news/stable-diffusion-3-research-paper</ext-link>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Saharia</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Fleet</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Norouzi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Salimans</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Cascaded diffusion models for high fidelity image generation</article-title>. <comment>arXiv:2106.15282. 2021 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2106.15282">http://arxiv.org/abs/2106.15282</ext-link>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Child</surname> <given-names>R</given-names></string-name>, <string-name><surname>Luan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Amodei</surname> <given-names>D</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Language models are unsupervised multitask learners [Internet]</article-title>. <publisher-name>OpenAI</publisher-name>; <year>2019 Feb 14 [cited 2025 Apr 8]</year>. <fpage>13</fpage> p. Available from: <ext-link ext-link-type="uri" xlink:href="https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf">https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ott</surname> <given-names>M</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Du</surname> <given-names>J</given-names></string-name>, <string-name><surname>Joshi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>RoBERTa: a robustly optimized BERT pretraining approach</article-title>. <comment>arXiv:1907.11692. 2019 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1907.11692">http://arxiv.org/abs/1907.11692</ext-link>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Raffel</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Matena</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Exploring the limits of transfer learning with a unified text-to-text transformer</article-title>. <source>J Mach Learn Res</source>. <year>2020</year>;<volume>21</volume>(<issue>140</issue>):<fpage>1</fpage>&#x2013;<lpage>67</lpage>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Brown</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mann</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ryder</surname> <given-names>N</given-names></string-name>, <string-name><surname>Subbiah</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kaplan</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Language models are few-shot learners</article-title>. In: <conf-name>Advances in Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2020 [cited 2025 Apr 8]</year>. p. <fpage>1877</fpage>&#x2013;<lpage>901</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html">https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html</ext-link>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Biderman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Black</surname> <given-names>S</given-names></string-name>, <string-name><surname>Golding</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hoppe</surname> <given-names>T</given-names></string-name>, <string-name><surname>Foster</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The Pile: an 800GB dataset of diverse text for language modeling</article-title>. <comment>arXiv:2101.00027. 2020 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2101.00027">http://arxiv.org/abs/2101.00027</ext-link>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Thoppilan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Freitas</surname> <given-names>DD</given-names></string-name>, <string-name><surname>Hall</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kulshreshtha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>HT</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LaMDA: language models for dialog applications</article-title>. <comment>arXiv:2201.08239. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2201.08239">http://arxiv.org/abs/2201.08239</ext-link>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chowdhery</surname> <given-names>A</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bosma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>G</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>PaLM: scaling language modeling with pathways</article-title>. <comment>arXiv:2204.02311. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2204.02311">http://arxiv.org/abs/2204.02311</ext-link>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>OpenAI</collab>, <string-name><surname>Achiam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Adler</surname> <given-names>S</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>L</given-names></string-name>, <string-name><surname>Akkaya</surname> <given-names>I</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GPT-4 technical report</article-title>. <comment>arXiv:2303.08774. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2303.08774">http://arxiv.org/abs/2303.08774</ext-link>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Touvron</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lavril</surname> <given-names>T</given-names></string-name>, <string-name><surname>Izacard</surname> <given-names>G</given-names></string-name>, <string-name><surname>Martinet</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lachaux</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Lacroix</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LLaMA: open and efficient foundation language models</article-title>. <comment>arXiv:2302.13971. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2302.13971">http://arxiv.org/abs/2302.13971</ext-link>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Attention is all you need</article-title>. In: <conf-name>Advances in Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017 [cited 2025 Apr 8]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html">https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html</ext-link>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Du</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>GLM: general language model pretraining with autoregressive blank infilling</article-title>. <comment>arXiv:2103.10360. 2022 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2103.10360">http://arxiv.org/abs/2103.10360</ext-link>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lynch</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Jensen</surname> <given-names>E</given-names></string-name>, <string-name><surname>Munro</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Zamponi</surname> <given-names>V</given-names></string-name>, <string-name><surname>Martinez</surname> <given-names>J</given-names></string-name>, <string-name><surname>O&#x2019;Brien</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>GPT-4 generated narratives of life events using a structured narrative prompt: a validation study</article-title>. <comment>arXiv:2402.05435. 2024 [cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2402.05435">http://arxiv.org/abs/2402.05435</ext-link>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Murtadha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>RoFormer: enhanced transformer with rotary position embedding</article-title>. <comment>arXiv:2104.09864. 2023 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2104.09864">http://arxiv.org/abs/2104.09864</ext-link>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Press O</collab>, <string-name><surname>Smith</surname> <given-names>NA</given-names></string-name>, <string-name><surname>Lewis</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Train short, test long: attention with linear biases enables input length extrapolation</article-title>. <comment>arXiv:2108.12409. 2022 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2108.12409">http://arxiv.org/abs/2108.12409</ext-link>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kitaev</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kaiser</surname> <given-names>&#x0141;.</given-names></string-name>, <string-name><surname>Levskaya</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Reformer: the efficient transformer</article-title>. <comment>arXiv:2001.04451. 2020 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2001.04451">http://arxiv.org/abs/2001.04451</ext-link>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zaheer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Guruganesh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Dubey</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ainslie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alberti</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ontanon</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Big bird: transformers for longer sequences</article-title>. <comment>arXiv:2007.14062. 2021 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2007.14062">http://arxiv.org/abs/2007.14062</ext-link>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>DY</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rudra</surname> <given-names>A</given-names></string-name>, <string-name><surname>R&#x00E9;</surname> <given-names>C</given-names></string-name></person-group>. <article-title>FlashAttention: fast and memory-efficient exact attention with IO-awareness</article-title>. <comment>arXiv:2205.14135. 2022 [cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2205.14135">http://arxiv.org/abs/2205.14135</ext-link>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>SQ</given-names></string-name></person-group>. <article-title>Parameter-efficient fine-tuning for large models: a comprehensive survey</article-title>. <comment>arXiv:2403.14608. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2403.14608">http://arxiv.org/abs/2403.14608</ext-link>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>EJ</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wallis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Allen-Zhu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LoRA: low-rank adaptation of large language models</article-title>. <comment>arXiv:2106.09685. 2021 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2106.09685">http://arxiv.org/abs/2106.09685</ext-link>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Houlsby</surname> <given-names>N</given-names></string-name>, <string-name><surname>Giurgiu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jastrzebski</surname> <given-names>S</given-names></string-name>, <string-name><surname>Morrone</surname> <given-names>B</given-names></string-name>, <string-name><surname>Laroussilhe</surname> <given-names>QD</given-names></string-name>, <string-name><surname>Gesmundo</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Parameter-efficient transfer learning for NLP</article-title>. In: <conf-name>Proceedings of the 36th International Conference on Machine Learning; 2019 Jun 9&#x2013;15</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>; <year>2019 [cited 2025 May 21]</year>. p. <fpage>2790</fpage>&#x2013;<lpage>9</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v97/houlsby19a.html">https://proceedings.mlr.press/v97/houlsby19a.html</ext-link>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>XL</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Prefix-tuning: optimizing continuous prompts for generation</article-title>. <comment>arXiv:2101.00190. 2021 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2101.00190">http://arxiv.org/abs/2101.00190</ext-link>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lester</surname> <given-names>B</given-names></string-name>, <string-name><surname>Al-Rfou</surname> <given-names>R</given-names></string-name>, <string-name><surname>Constant</surname> <given-names>N</given-names></string-name></person-group>. <article-title>The power of scale for parameter-efficient prompt tuning</article-title>. <comment>arXiv:2104.08691. 2021 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2104.08691">http://arxiv.org/abs/2104.08691</ext-link>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Arora</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Geiger</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jurafsky</surname> <given-names>D</given-names></string-name>, <string-name><surname>Manning</surname> <given-names>CD</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ReFT: representation finetuning for language models</article-title>. <comment>arXiv:2404.03592. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2404.03592">http://arxiv.org/abs/2404.03592</ext-link>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Agrawala</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Adding conditional control to text-to-image diffusion models</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2023 Oct 1&#x2013;6</conf-name>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>3813</fpage>&#x2013;<lpage>24</lpage>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ruiz</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jampani</surname> <given-names>V</given-names></string-name>, <string-name><surname>Pritch</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Rubinstein</surname> <given-names>M</given-names></string-name>, <string-name><surname>Aberman</surname> <given-names>K</given-names></string-name></person-group>. <article-title>DreamBooth: fine tuning text-to-image diffusion models for subject-driven generation</article-title>. <comment>arXiv:2208.12242. 2023 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2208.12242">http://arxiv.org/abs/2208.12242</ext-link>.</mixed-citation></ref>
<ref id="ref-80"><label>[80]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gal</surname> <given-names>R</given-names></string-name>, <string-name><surname>Alaluf</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Atzmon</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Patashnik</surname> <given-names>O</given-names></string-name>, <string-name><surname>Bermano</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Chechik</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>An image is worth one word: personalizing text-to-image generation using textual inversion</article-title>. <comment>arXiv:2208.01618. 2022 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2208.01618">http://arxiv.org/abs/2208.01618</ext-link>.</mixed-citation></ref>
<ref id="ref-81"><label>[81]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ha</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>A</given-names></string-name>, <string-name><surname>Le</surname> <given-names>QV</given-names></string-name></person-group>. <article-title>HyperNetworks</article-title>. <comment>arXiv:1609.09106. 2016 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1609.09106">http://arxiv.org/abs/1609.09106</ext-link>.</mixed-citation></ref>
<ref id="ref-82"><label>[82]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Davila</surname> <given-names>A</given-names></string-name>, <string-name><surname>Colan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hasegawa</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Comparison of fine-tuning strategies for transfer learning in medical image classification</article-title>. <source>Image Vis Comput</source>. <year>2024</year>;<volume>146</volume>(<issue>2</issue>):<fpage>105012</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.imavis.2024.105012</pub-id>.</mixed-citation></ref>
<ref id="ref-83"><label>[83]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Parthasarathy</surname> <given-names>VB</given-names></string-name>, <string-name><surname>Zafar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shahid</surname> <given-names>A</given-names></string-name></person-group>. <article-title>The ultimate guide to fine-tuning LLMs from basics to breakthroughs: an exhaustive review of technologies, research, best practices, applied research challenges and opportunities</article-title>. <comment>arXiv:2408.13296. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2408.13296">http://arxiv.org/abs/2408.13296</ext-link>.</mixed-citation></ref>
<ref id="ref-84"><label>[84]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lippmann</surname> <given-names>P</given-names></string-name>, <string-name><surname>Skublicki</surname> <given-names>K</given-names></string-name>, <string-name><surname>Tanner</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ishiwatari</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Context-Informed machine translation of manga using multimodal large language models</article-title>. <comment>arXiv:2411.02589. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.02589">http://arxiv.org/abs/2411.02589</ext-link>.</mixed-citation></ref>
<ref id="ref-85"><label>[85]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sachdeva</surname> <given-names>R</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Tails tell tales: chapter-wide manga transcriptions with character names</article-title>. <comment>arXiv:2408.00298. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2408.00298">http://arxiv.org/abs/2408.00298</ext-link>.</mixed-citation></ref>
<ref id="ref-86"><label>[86]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Vivoli</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bertini</surname> <given-names>M</given-names></string-name>, <string-name><surname>Karatzas</surname> <given-names>D</given-names></string-name></person-group>. <article-title>CoMix: a comprehensive benchmark for multi-task comic understanding</article-title>. <comment>arXiv:2407.03550. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2407.03550">http://arxiv.org/abs/2407.03550</ext-link>.</mixed-citation></ref>
<ref id="ref-87"><label>[87]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ju</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Human-Art: a versatile human-centric dataset bridging natural and artificial scenes</article-title>. <comment>arXiv:2303.02760. 2023 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2303.02760">http://arxiv.org/abs/2303.02760</ext-link>.</mixed-citation></ref>
<ref id="ref-88"><label>[88]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Aizawa</surname> <given-names>K</given-names></string-name>, <string-name><surname>Matsui</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Manga109Dialog: a large-scale dialogue dataset for comics speaker detection</article-title>. <comment>arXiv:2306.17469. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2306.17469">http://arxiv.org/abs/2306.17469</ext-link>.</mixed-citation></ref>
<ref id="ref-89"><label>[89]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Seo</surname> <given-names>CW</given-names></string-name>, <string-name><surname>Ashtari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Noh</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Semi-supervised reference-based sketch extraction using a contrastive learning framework</article-title>. <comment>arXiv:2407.14026. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2407.14026">http://arxiv.org/abs/2407.14026</ext-link>.</mixed-citation></ref>
<ref id="ref-90"><label>[90]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Parsing-conditioned anime translation: a new dataset and method</article-title>. <source>ACM Trans Graph</source>. <year>2023</year>;<volume>42</volume>(<issue>3</issue>):<fpage>30:1</fpage>&#x2013;<lpage>30:14</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3585002</pub-id>.</mixed-citation></ref>
<ref id="ref-91"><label>[91]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Siyao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Loy</surname> <given-names>CC</given-names></string-name></person-group>. <article-title>AnimeRun: 2D animation visual correspondence from open source 3D movies</article-title>. <comment>arXiv:2211.05709. 2022 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2211.05709">http://arxiv.org/abs/2211.05709</ext-link>.</mixed-citation></ref>
<ref id="ref-92"><label>[92]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Baek</surname> <given-names>J</given-names></string-name>, <string-name><surname>Matsui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Aizawa</surname> <given-names>K</given-names></string-name></person-group>. <article-title>COO: comic onomatopoeia dataset for recognizing arbitrary or truncated texts</article-title>. <comment>arXiv:2207.04675. 2022 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2207.04675">http://arxiv.org/abs/2207.04675</ext-link>.</mixed-citation></ref>
<ref id="ref-93"><label>[93]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>K</given-names></string-name>, <string-name><surname>Park</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>AnimeCeleb: large-scale animation celebheads dataset for head reenactment</article-title>. <comment>arXiv:2111.07640. 2022 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2111.07640">http://arxiv.org/abs/2111.07640</ext-link>.</mixed-citation></ref>
<ref id="ref-94"><label>[94]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Rios</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>WH</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>BC</given-names></string-name></person-group>. <article-title>DAF: re: a challenging, crowd-sourced, large-scale, long-tailed dataset for anime character recognition</article-title>. <comment>arXiv:2101.08674. 2021 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2101.08674">http://arxiv.org/abs/2101.08674</ext-link>.</mixed-citation></ref>
<ref id="ref-95"><label>[95]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Cartoon face recognition: a benchmark dataset</article-title>. <comment>arXiv:1907.13394. 2020 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1907.13394">http://arxiv.org/abs/1907.13394</ext-link>.</mixed-citation></ref>
<ref id="ref-96"><label>[96]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gobbo</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Herrera</surname> <given-names>RM</given-names></string-name></person-group>. <article-title>Unconstrained text detection in manga: a new dataset and baseline</article-title>. <comment>arXiv:2009.04042. 2020 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2009.04042">http://arxiv.org/abs/2009.04042</ext-link>.</mixed-citation></ref>
<ref id="ref-97"><label>[97]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name></person-group>. <chapter-title>An illustration region dataset</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Vedaldi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bischof</surname> <given-names>H</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name>, <string-name><surname>Frahm</surname> <given-names>JM</given-names></string-name></person-group>, editors. <source>Computer vision&#x2014;ECCV 2020, Lecture notes in computer science</source>. Vol. <volume>12358</volume>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>; <year>2020 [cited 2025 May 21]</year>. p. <fpage>137</fpage>&#x2013;<lpage>54</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://link.springer.com/10.1007/978-3-030-58601-0_9">https://link.springer.com/10.1007/978-3-030-58601-0_9</ext-link>.</mixed-citation></ref>
<ref id="ref-98"><label>[98]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aizawa</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fujimoto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Otsubo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ogawa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Matsui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tsubota</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Building a manga dataset &#x201C;Manga109&#x201D; with annotations for multimedia applications</article-title>. <source>IEEE Multimed</source>. <year>2020</year>;<volume>27</volume>(<issue>2</issue>):<fpage>8</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mmul.2020.2987895</pub-id>.</mixed-citation></ref>
<ref id="ref-99"><label>[99]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Karras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Laine</surname> <given-names>S</given-names></string-name>, <string-name><surname>Aittala</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hellsten</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lehtinen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Aila</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Analyzing and improving the image quality of StyleGAN</article-title>. <comment>arXiv:1912.04958. 2020 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1912.04958">http://arxiv.org/abs/1912.04958</ext-link>.</mixed-citation></ref>
<ref id="ref-100"><label>[100]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A survey of image style transfer research</article-title>. In: <conf-name>2022 2nd International Conference on Computer Science, Electronic Information Engineering and Intelligent Control Technology (CEI); 2022 Sep 23&#x2013;25</conf-name>; <publisher-loc>Nanjing, China</publisher-loc>; <year>2022 [cited 2025 May 21]</year>. p. <fpage>133</fpage>&#x2013;<lpage>7</lpage>. Available from: <ext-link ext-link-type="uri" xlink:href="https://ieeexplore.ieee.org/document/9950226">https://ieeexplore.ieee.org/document/9950226</ext-link>.</mixed-citation></ref>
<ref id="ref-101"><label>[101]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Multimodal alignment and fusion: a survey</article-title>. <comment>arXiv:2411.17040. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.17040">http://arxiv.org/abs/2411.17040</ext-link>.</mixed-citation></ref>
<ref id="ref-102"><label>[102]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Olamendy</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>Tackling the challenge of imbalanced datasets: a comprehensive guide. Medium</article-title>. <year>2024</year>. <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://medium.com/@juanc.olamendy/tackling-the-challenge-of-imbalanced-datasets-a-comprehensive-guide-2feb11ca2fa0">https://medium.com/@juanc.olamendy/tackling-the-challenge-of-imbalanced-datasets-a-comprehensive-guide-2feb11ca2fa0</ext-link>.</mixed-citation></ref>
<ref id="ref-103"><label>[103]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>black-forest-labs/flux</collab></person-group>. <article-title>black-forest-labs</article-title>; <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/black-forest-labs/flux">https://github.com/black-forest-labs/flux</ext-link>.</mixed-citation></ref>
<ref id="ref-104"><label>[104]</label><mixed-citation publication-type="other"><article-title>Midjourney</article-title>; <year>2025</year> <comment>[cited 2025 Jul 20]. Updates</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.midjourney.com/website">https://www.midjourney.com/website</ext-link>.</mixed-citation></ref>
<ref id="ref-105"><label>[105]</label><mixed-citation publication-type="other"><article-title>OpenAI Help Center</article-title>. [cited 2025 Jul 20]. <comment>ChatGPT&#x2014;Release Notes</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://help.openai.com/en/articles/6825453-chatgpt-release-notes#h_8670d7da97">https://help.openai.com/en/articles/6825453-chatgpt-release-notes#h_8670d7da97</ext-link>.</mixed-citation></ref>
<ref id="ref-106"><label>[106]</label><mixed-citation publication-type="other"><article-title>Adobe Firefly-Free Generative AI for creatives</article-title>. <comment>[cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.adobe.com/products/firefly.html">https://www.adobe.com/products/firefly.html</ext-link>.</mixed-citation></ref>
<ref id="ref-107"><label>[107]</label><mixed-citation publication-type="other"><article-title>Ideogram 3.0</article-title>. <comment>[cited 2025 Jul 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://about.ideogram.ai/3.0">https://about.ideogram.ai/3.0</ext-link>.</mixed-citation></ref>
<ref id="ref-108"><label>[108]</label><mixed-citation publication-type="other"><article-title>THUDM/CogView4. Z.ai &#x0026; THUKEG</article-title>; <year>2025</year> <comment>[cited 2025 Jul 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/THUDM/CogView4">https://github.com/THUDM/CogView4</ext-link>.</mixed-citation></ref>
<ref id="ref-109"><label>[109]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Kwai-Kolors/Kolors</collab></person-group>. <article-title>Kolors</article-title>; <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/Kwai-Kolors/Kolors">https://github.com/Kwai-Kolors/Kolors</ext-link>.</mixed-citation></ref>
<ref id="ref-110"><label>[110]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Seedream 3.0 technical report</article-title>. <comment>arXiv:2504.11346. 2025 [cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2504.11346">http://arxiv.org/abs/2504.11346</ext-link>.</mixed-citation></ref>
<ref id="ref-111"><label>[111]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Long</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hunyuan-DiT: a powerful multi-resolution diffusion transformer with fine-grained chinese understanding</article-title>. <comment>arXiv:2405.08748. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2405.08748">http://arxiv.org/abs/2405.08748</ext-link>.</mixed-citation></ref>
<ref id="ref-112"><label>[112]</label><mixed-citation publication-type="other"><article-title>Flux AI-Free Online Advanced Flux AI Image Generator</article-title>. <comment>[cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://flux1.ai/">https://flux1.ai/</ext-link>.</mixed-citation></ref>
<ref id="ref-113"><label>[113]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Mou</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>T2I-Adapter: learning adapters to dig out more controllable ability for text-to-image diffusion models</article-title>. <comment>arXiv:2302.08453. 2023 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2302.08453">http://arxiv.org/abs/2302.08453</ext-link>.</mixed-citation></ref>
<ref id="ref-114"><label>[114]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kawano</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Aoki</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>MaskDiffusion: exploiting pre-trained diffusion models for semantic segmentation</article-title>. <comment>arXiv:2403.11194. 2024 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2403.11194">http://arxiv.org/abs/2403.11194</ext-link>.</mixed-citation></ref>
<ref id="ref-115"><label>[115]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Latent consistency models: synthesizing high-resolution images with few-step inference</article-title>. <comment>arXiv:2310.04378. 2023 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2310.04378">http://arxiv.org/abs/2310.04378</ext-link>.</mixed-citation></ref>
<ref id="ref-116"><label>[116]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>AUTOMATIC1111</collab></person-group>. <article-title>AUTOMATIC1111/stable-diffusion-webui-feature-showcase</article-title>. <year>2025</year> <comment>[cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/AUTOMATIC1111/stable-diffusion-webui-feature-showcase">https://github.com/AUTOMATIC1111/stable-diffusion-webui-feature-showcase</ext-link>.</mixed-citation></ref>
<ref id="ref-117"><label>[117]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>comfyanonymous</collab></person-group>. <article-title>comfyanonymous/ComfyUI</article-title>. <year>2025</year> <comment>[cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/comfyanonymous/ComfyUI">https://github.com/comfyanonymous/ComfyUI</ext-link>.</mixed-citation></ref>
<ref id="ref-118"><label>[118]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>lllyasviel</collab></person-group>. <article-title>lllyasviel/Paints-UNDO</article-title>. <year>2025</year> <comment>[cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/lllyasviel/Paints-UNDO">https://github.com/lllyasviel/Paints-UNDO</ext-link>.</mixed-citation></ref>
<ref id="ref-119"><label>[119]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Imagen-Team-Google</collab>, <string-name><surname>Baldridge</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bauer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bhutani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Brichtova</surname> <given-names>N</given-names></string-name>, <string-name><surname>Bunner</surname> <given-names>A</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Imagen 3</article-title>. <comment>arXiv:2408.07009. 2024 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2408.07009">http://arxiv.org/abs/2408.07009</ext-link>.</mixed-citation></ref>
<ref id="ref-120"><label>[120]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Team</surname> <given-names>G</given-names></string-name>, <string-name><surname>Mesnard</surname> <given-names>T</given-names></string-name>, <string-name><surname>Hardin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Dadashi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bhupatiraju</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pathak</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Gemma: open models based on gemini research and technology</article-title>. <comment>arXiv:2403.08295. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2403.08295">http://arxiv.org/abs/2403.08295</ext-link>.</mixed-citation></ref>
<ref id="ref-121"><label>[121]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhuo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Lumina-Image 2.0: a unified and efficient image generative framework</article-title>. <comment>arXiv:2503.21758. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.21758">http://arxiv.org/abs/2503.21758</ext-link>.</mixed-citation></ref>
<ref id="ref-122"><label>[122]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Silver</surname> <given-names>D</given-names></string-name>, <string-name><surname>Lever</surname> <given-names>G</given-names></string-name>, <string-name><surname>Heess</surname> <given-names>N</given-names></string-name>, <string-name><surname>Degris</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wierstra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Riedmiller</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Deterministic policy gradient algorithms</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Machine Learning</conf-name>. <publisher-loc>Bejing, China</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2014</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v32/silver14.html">https://proceedings.mlr.press/v32/silver14.html</ext-link>.</mixed-citation></ref>
<ref id="ref-123"><label>[123]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ghosh</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hajishirzi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>L</given-names></string-name></person-group>. <article-title>GenEval: an object-focused framework for evaluating text-to-image alignment</article-title>. <comment>arXiv:2310.11513. 2023 [cited 2025 Apr 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2310.11513">http://arxiv.org/abs/2310.11513</ext-link>.</mixed-citation></ref>
<ref id="ref-124"><label>[124]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Teng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CogView3: finer and faster text-to-image generation via relay diffusion</article-title>. <comment>arXiv:2403.05121. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2403.05121">http://arxiv.org/abs/2403.05121</ext-link>.</mixed-citation></ref>
<ref id="ref-125"><label>[125]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Betker</surname> <given-names>J</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Jing</surname> <given-names>L</given-names></string-name>, <string-name><surname>Brooks</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Improving image generation with better captions [Internet]</article-title>. <publisher-name>OpenAI</publisher-name>; <year>2023</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://cdn.openai.com/papers/dall-e-3.pdf">https://cdn.openai.com/papers/dall-e-3.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-126"><label>[126]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Koh</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>D</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Illustrious: an open advanced illustration model</article-title>. <comment>arXiv:2409.19946. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2409.19946">http://arxiv.org/abs/2409.19946</ext-link>.</mixed-citation></ref>
<ref id="ref-127"><label>[127]</label><mixed-citation publication-type="other"><article-title>Pony Diffusion V6 XL - V6 (start with this one) | Stable Diffusion Checkpoint | Civitai</article-title>. <year>2025</year> <comment>[cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://civitai.com/models/257749/pony-diffusion-v6-xl">https://civitai.com/models/257749/pony-diffusion-v6-xl</ext-link>.</mixed-citation></ref>
<ref id="ref-128"><label>[128]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sehwag</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Spranger</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Stretching each dollar: diffusion training from scratch on a micro-budget</article-title>. <comment>arXiv:2407.15811. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2407.15811">http://arxiv.org/abs/2407.15811</ext-link>.</mixed-citation></ref>
<ref id="ref-129"><label>[129]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>E</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SANA: efficient high-resolution image synthesis with linear diffusion transformers</article-title>. <comment>arXiv:2410.10629. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.10629">http://arxiv.org/abs/2410.10629</ext-link>.</mixed-citation></ref>
<ref id="ref-130"><label>[130]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Mi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>KC</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Tulyakov</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>I think, therefore i diffuse: enabling multimodal in-context reasoning in diffusion models</article-title>. <comment>arXiv:2502.10458. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.10458">http://arxiv.org/abs/2502.10458</ext-link>.</mixed-citation></ref>
<ref id="ref-131"><label>[131]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Vinci</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multimodal representation alignment for image generation: text-image interleaved control is easier than you think</article-title>. <comment>arXiv:2502.20172. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.20172">http://arxiv.org/abs/2502.20172</ext-link>.</mixed-citation></ref>
<ref id="ref-132"><label>[132]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ouyang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MangaNinja: line art colorization with precise reference following</article-title>. <comment>arXiv:2501.08332. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.08332">http://arxiv.org/abs/2501.08332</ext-link>.</mixed-citation></ref>
<ref id="ref-133"><label>[133]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shou</surname> <given-names>MZ</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>PhotoDoodle: learning artistic image editing from few-shot pairwise data</article-title>. <comment>arXiv:2502.14397. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.14397">http://arxiv.org/abs/2502.14397</ext-link>.</mixed-citation></ref>
<ref id="ref-134"><label>[134]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>One diffusion step to real-world super-resolution via flow trajectory distillation</article-title>. <comment>arXiv:2502.01993. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.01993">http://arxiv.org/abs/2502.01993</ext-link>.</mixed-citation></ref>
<ref id="ref-135"><label>[135]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Content-style decoupling for unsupervised makeup transfer without generating pseudo ground truth</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>7601</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-136"><label>[136]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ye</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>IP-Adapter: text compatible image prompt adapter for text-to-image diffusion models</article-title>. <comment>arXiv:2308.06721. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2308.06721">http://arxiv.org/abs/2308.06721</ext-link>.</mixed-citation></ref>
<ref id="ref-137"><label>[137]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>InstantID: zero-shot identity-preserving generation in seconds</article-title>. <comment>arXiv:2401.07519. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2401.07519">http://arxiv.org/abs/2401.07519</ext-link>.</mixed-citation></ref>
<ref id="ref-138"><label>[138]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Qi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>PhotoMaker: customizing realistic human photos via stacked ID embedding</article-title>. <comment>arXiv:2312.04461. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2312.04461">http://arxiv.org/abs/2312.04461</ext-link>.</mixed-citation></ref>
<ref id="ref-139"><label>[139]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>StoryDiffusion: consistent self-attention for long-range image and video generation</article-title>. <source>Adv Neural Inf Process Syst</source>. <year>2024</year>;<volume>37</volume>:<fpage>110315</fpage>&#x2013;<lpage>40</lpage>.</mixed-citation></ref>
<ref id="ref-140"><label>[140]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tong</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>DiffSensei: bridging multi-modal LLMs and diffusion models for customized manga generation</article-title>. <comment>arXiv:2412.07589. 2025 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.07589">http://arxiv.org/abs/2412.07589</ext-link>.</mixed-citation></ref>
<ref id="ref-141"><label>[141]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tuo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>C</given-names></string-name>, <string-name><surname>Geng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bo</surname> <given-names>L</given-names></string-name></person-group>. <article-title>AnyStory: towards unified single and multiple subject personalization in text-to-image generation</article-title>. <comment>arXiv:2501.09503. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.09503">http://arxiv.org/abs/2501.09503</ext-link>.</mixed-citation></ref>
<ref id="ref-142"><label>[142]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Brock</surname> <given-names>A</given-names></string-name>, <string-name><surname>Donahue</surname> <given-names>J</given-names></string-name>, <string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Large scale GAN training for high fidelity natural image synthesis</article-title>. <comment>arXiv:1809.11096. 2019 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1809.11096">http://arxiv.org/abs/1809.11096</ext-link>.</mixed-citation></ref>
<ref id="ref-143"><label>[143]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bovik</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Sheikh</surname> <given-names>HR</given-names></string-name>, <string-name><surname>Simoncelli</surname> <given-names>EP</given-names></string-name></person-group>. <article-title>Image quality assessment: from error visibility to structural similarity</article-title>. <source>IEEE Trans Image Process</source>. <year>2004</year>;<volume>13</volume>(<issue>4</issue>):<fpage>600</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tip.2003.819861</pub-id>; <pub-id pub-id-type="pmid">15376593</pub-id></mixed-citation></ref>
<ref id="ref-144"><label>[144]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Salimans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name>, <string-name><surname>Zaremba</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cheung</surname> <given-names>V</given-names></string-name>, <string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Improved techniques for training GANs</article-title>. <comment>arXiv:1606.03498. 2016 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1606.03498">http://arxiv.org/abs/1606.03498</ext-link>.</mixed-citation></ref>
<ref id="ref-145"><label>[145]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Isola</surname> <given-names>P</given-names></string-name>, <string-name><surname>Efros</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Shechtman</surname> <given-names>E</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>O</given-names></string-name></person-group>. <article-title>The unreasonable effectiveness of deep features as a perceptual metric</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2018 Jun 20&#x2013;23</conf-name>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>. p. <fpage>586</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-146"><label>[146]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bi&#x0144;kowski</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sutherland</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Arbel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gretton</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Demystifying MMD GANs</article-title>. <comment>arXiv:1801.01401. 2021 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1801.01401">http://arxiv.org/abs/1801.01401</ext-link>.</mixed-citation></ref>
<ref id="ref-147"><label>[147]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Sora: a review on background, technology, limitations, and opportunities of large vision models</article-title>. <comment>arXiv:2402.17177. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2402.17177">http://arxiv.org/abs/2402.17177</ext-link>.</mixed-citation></ref>
<ref id="ref-148"><label>[148]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>team TVPU</collab></person-group>. <article-title>Veo product updates</article-title>. <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://updates.veo.co/">https://updates.veo.co/</ext-link>.</mixed-citation></ref>
<ref id="ref-149"><label>[149]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Polyak</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zohar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Brown</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tjandra</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sinha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>A</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Movie gen: a cast of media foundation models</article-title>. <comment>arXiv:2410.13720. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.13720">http://arxiv.org/abs/2410.13720</ext-link>.</mixed-citation></ref>
<ref id="ref-150"><label>[150]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>HaCohen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chiprut</surname> <given-names>N</given-names></string-name>, <string-name><surname>Brazowski</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shalem</surname> <given-names>D</given-names></string-name>, <string-name><surname>Moshe</surname> <given-names>D</given-names></string-name>, <string-name><surname>Richardson</surname> <given-names>E</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LTX-Video: realtime video latent diffusion</article-title>. <comment>arXiv:2501.00103. 2024 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.00103">http://arxiv.org/abs/2501.00103</ext-link>.</mixed-citation></ref>
<ref id="ref-151"><label>[151]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>genmoai/mochi</collab></person-group>. <article-title>Genmo</article-title>; <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/genmoai/mochi">https://github.com/genmoai/mochi</ext-link>.</mixed-citation></ref>
<ref id="ref-152"><label>[152]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>CogVideo: large-scale pretraining for text-to-video generation via transformers</article-title>. <comment>arXiv:2205.15868. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2205.15868">http://arxiv.org/abs/2205.15868</ext-link>.</mixed-citation></ref>
<ref id="ref-153"><label>[153]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Min</surname> <given-names>R</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>HunyuanVideo: a systematic framework for large video generative models</article-title>. <comment>arXiv:2412.03603. 2025 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.03603">http://arxiv.org/abs/2412.03603</ext-link>.</mixed-citation></ref>
<ref id="ref-154"><label>[154]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>NNVIDIA</collab>, <string-name><surname>Agarwal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bala</surname> <given-names>M</given-names></string-name>, <string-name><surname>Balaji</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Barker</surname> <given-names>E</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Cosmos world foundation model platform for physical AI</article-title>. <comment>arXiv:2501.03575. 2025 [cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.03575">http://arxiv.org/abs/2501.03575</ext-link>.</mixed-citation></ref>
<ref id="ref-155"><label>[155]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qiu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Fei</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SkyReels-A1: expressive portrait animation in video diffusion transformers</article-title>. <comment>arXiv:2502.10841. 2025 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.10841">http://arxiv.org/abs/2502.10841</ext-link>.</mixed-citation></ref>
<ref id="ref-156"><label>[156]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>WanTeam</collab>, <string-name><surname>Wang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ai</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>CW</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Wan: open and advanced large-scale video generative models</article-title>. <comment>arXiv:2503.20314. 2025 [cited 2025 Apr 18]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.20314">http://arxiv.org/abs/2503.20314</ext-link>.</mixed-citation></ref>
<ref id="ref-157"><label>[157]</label><mixed-citation publication-type="other"><article-title>Runway | Tools for human imagination</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://runwayml.com/changelog">https://runwayml.com/changelog</ext-link>.</mixed-citation></ref>
<ref id="ref-158"><label>[158]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>pikalabsorg</collab></person-group>. <article-title>Pika 2.2. Pika Labs</article-title>. <year>2025 [cited 2025 Jul 21]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://pikalabs.org/pika-2-2/">https://pikalabs.org/pika-2-2/</ext-link>.</mixed-citation></ref>
<ref id="ref-159"><label>[159]</label><mixed-citation publication-type="other"><article-title>Luma AI | AI Video Generation with Ray2 &#x0026; Dream Machine | Luma AI</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://lumalabs.ai/">https://lumalabs.ai/</ext-link>.</mixed-citation></ref>
<ref id="ref-160"><label>[160]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Blattmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dockhorn</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kulal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mendelevitch</surname> <given-names>D</given-names></string-name>, <string-name><surname>Kilian</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lorenz</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Stable video diffusion: scaling latent video diffusion models to large datasets</article-title>. <comment>arXiv:2311.15127. 2023 [cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2311.15127">http://arxiv.org/abs/2311.15127</ext-link>.</mixed-citation></ref>
<ref id="ref-161"><label>[161]</label><mixed-citation publication-type="other"><article-title>Kling AI: Next-Gen AI Video &#x0026; AI Image Generator</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://app.klingai.com/global/">https://app.klingai.com/global/</ext-link>.</mixed-citation></ref>
<ref id="ref-162"><label>[162]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>hailuoai.video</collab></person-group>. <article-title>Hailuo AI: Transform Idea to Visual with AI&#x2013;AI Generated Video | Hailuo AI</article-title>. <comment>[cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://hailuoai.video">https://hailuoai.video</ext-link>.</mixed-citation></ref>
<ref id="ref-163"><label>[163]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>N</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Pyramidal flow matching for efficient video generative modeling</article-title>. <comment>arXiv:2410.05954. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.05954">http://arxiv.org/abs/2410.05954</ext-link>.</mixed-citation></ref>
<ref id="ref-164"><label>[164]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>PyramidFlow: high-resolution defect contrastive localization using pyramid normalizing flow</article-title>. In: <conf-name>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>14143</fpage>&#x2013;<lpage>52</lpage>.</mixed-citation></ref>
<ref id="ref-165"><label>[165]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ziqiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>C</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pi</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VideoDPO: omni-preference alignment for video diffusion generation</article-title>. <comment>arXiv:2412.14167. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.14167">http://arxiv.org/abs/2412.14167</ext-link>.</mixed-citation></ref>
<ref id="ref-166"><label>[166]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>JZ</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ran</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Show-1: marrying pixel and latent diffusion models for text-to-video generation</article-title>. <source>Int J Comput Vis</source>. <year>2025</year>;<volume>133</volume>(<issue>4</issue>):<fpage>1879</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-024-02271-9</pub-id>.</mixed-citation></ref>
<ref id="ref-167"><label>[167]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>VideoCrafter2: overcoming data limitations for high-quality video diffusion models</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>7310</fpage>&#x2013;<lpage>20</lpage>.</mixed-citation></ref>
<ref id="ref-168"><label>[168]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Owl-1: omni world model for consistent long video generation</article-title>. <comment>arXiv:2412.09600. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.09600">http://arxiv.org/abs/2412.09600</ext-link>.</mixed-citation></ref>
<ref id="ref-169"><label>[169]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Siarohin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Menapace</surname> <given-names>W</given-names></string-name>, <string-name><surname>Skorokhodov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chordia</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mind the time: temporally-controlled multi-event video generation</article-title>. <comment>arXiv:2412.05263. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.05263">http://arxiv.org/abs/2412.05263</ext-link>.</mixed-citation></ref>
<ref id="ref-170"><label>[170]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ni</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>NUWA-XL: diffusion over diffusion for eXtremely long video generation</article-title>. <comment>arXiv:2303.12346. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2303.12346">http://arxiv.org/abs/2303.12346</ext-link>.</mixed-citation></ref>
<ref id="ref-171"><label>[171]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Long video diffusion generation with segmented cross-attention and content-rich video data curation</article-title>. <comment>arXiv:2412.01316. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.01316">http://arxiv.org/abs/2412.01316</ext-link>.</mixed-citation></ref>
<ref id="ref-172"><label>[172]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Atzmon</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gal</surname> <given-names>R</given-names></string-name>, <string-name><surname>Tewel</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kasten</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chechik</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Motion by queries: identity-motion trade-offs in text-to-video generation</article-title>. <comment>arXiv:2412.07750. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.07750">http://arxiv.org/abs/2412.07750</ext-link>.</mixed-citation></ref>
<ref id="ref-173"><label>[173]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liew</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MagicAnimate: temporally consistent human image animation using diffusion model</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>1481</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-174"><label>[174]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>AnimateAnything: consistent and controllable animation for video generation</article-title>. <comment>arXiv:2411.10836. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.10836">http://arxiv.org/abs/2411.10836</ext-link>.</mixed-citation></ref>
<ref id="ref-175"><label>[175]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gal</surname> <given-names>R</given-names></string-name>, <string-name><surname>Vinker</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Alaluf</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bermano</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cohen-Or</surname> <given-names>D</given-names></string-name>, <string-name><surname>Shamir</surname> <given-names>A</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Breathing life into sketches using text-to-video priors</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>4325</fpage>&#x2013;<lpage>36</lpage>.</mixed-citation></ref>
<ref id="ref-176"><label>[176]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>StyleCrafter: enhancing stylized text-to-video generation with style adapter</article-title>. <comment>arXiv:2312.00330. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2312.00330">http://arxiv.org/abs/2312.00330</ext-link>.</mixed-citation></ref>
<ref id="ref-177"><label>[177]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Reconstruction vs. generation: taming optimization dilemma in latent diffusion models</article-title>. <comment>arXiv:2501.01423. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.01423">http://arxiv.org/abs/2501.01423</ext-link>.</mixed-citation></ref>
<ref id="ref-178"><label>[178]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CineMaster: a 3D-aware and controllable framework for cinematic text-to-video generation</article-title>. <comment>arXiv:2502.08639. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.08639">http://arxiv.org/abs/2502.08639</ext-link>.</mixed-citation></ref>
<ref id="ref-179"><label>[179]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hyung</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Spatiotemporal skip guidance for enhanced video diffusion sampling</article-title>. <comment>arXiv:2411.18664. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.18664">http://arxiv.org/abs/2411.18664</ext-link>.</mixed-citation></ref>
<ref id="ref-180"><label>[180]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>He</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Phantom: subject-consistent video generation via cross-modal alignment</article-title>. <comment>arXiv:2502.11079. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.11079">http://arxiv.org/abs/2502.11079</ext-link>.</mixed-citation></ref>
<ref id="ref-181"><label>[181]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhuang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qiao</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Vlogger: make your dream a vlog</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>8806</fpage>&#x2013;<lpage>17</lpage>.</mixed-citation></ref>
<ref id="ref-182"><label>[182]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mei</surname> <given-names>T</given-names></string-name></person-group>. <chapter-title>VideoStudio: generating consistent-content and multi-scene videos</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>468</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-73027-6_27</pub-id>.</mixed-citation></ref>
<ref id="ref-183"><label>[183]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>A</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VFX creator: animated visual effect generation with controllable diffusion transformer</article-title>. <comment>arXiv:2502.05979. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.05979">http://arxiv.org/abs/2502.05979</ext-link>.</mixed-citation></ref>
<ref id="ref-184"><label>[184]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Xing</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>DynamiCrafter: animating open-domain images with video diffusion priors</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>399</fpage>&#x2013;<lpage>417</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-72952-2_23</pub-id>.</mixed-citation></ref>
<ref id="ref-185"><label>[185]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>K</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>I2VGen-XL: high-quality image-to-video synthesis via cascaded diffusion models</article-title>. <comment>arXiv:2311.04145. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2311.04145">http://arxiv.org/abs/2311.04145</ext-link>.</mixed-citation></ref>
<ref id="ref-186"><label>[186]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cui</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hallo3: highly dynamic and realistic portrait image animation with video diffusion transformer</article-title>. <comment>arXiv:2412.00733. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.00733">http://arxiv.org/abs/2412.00733</ext-link>.</mixed-citation></ref>
<ref id="ref-187"><label>[187]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VQTalker: towards multilingual talking avatars through facial motion tokenization</article-title>. <comment>arXiv:2412.09892. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.09892">http://arxiv.org/abs/2412.09892</ext-link>.</mixed-citation></ref>
<ref id="ref-188"><label>[188]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MEMO: memory-guided diffusion for expressive talking video generation</article-title>. <comment>arXiv:2412.04448. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.04448">http://arxiv.org/abs/2412.04448</ext-link>.</mixed-citation></ref>
<ref id="ref-189"><label>[189]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LatentSync: taming audio-conditioned latent diffusion models for lip sync with SyncNet supervision</article-title>. <comment>arXiv:2412.09262. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.09262">http://arxiv.org/abs/2412.09262</ext-link>.</mixed-citation></ref>
<ref id="ref-190"><label>[190]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Rong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>INFP: audio-driven interactive head generation in dyadic conversations</article-title>. <comment>arXiv:2412.04037. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.04037">http://arxiv.org/abs/2412.04037</ext-link>.</mixed-citation></ref>
<ref id="ref-191"><label>[191]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cui</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Shang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hallo2: long-duration and high-resolution audio-driven portrait image animation</article-title>. <comment>arXiv:2410.07718. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.07718">http://arxiv.org/abs/2410.07718</ext-link>.</mixed-citation></ref>
<ref id="ref-192"><label>[192]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bo</surname> <given-names>L</given-names></string-name></person-group>. <article-title>EMO2: end-effector guided audio-driven avatar video generation</article-title>. <comment>arXiv:2501.10687. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.10687">http://arxiv.org/abs/2501.10687</ext-link>.</mixed-citation></ref>
<ref id="ref-193"><label>[193]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>IDOL: unified dual-modal latent diffusion for human-centric joint video-depth generation</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>134</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-72633-0_8</pub-id>.</mixed-citation></ref>
<ref id="ref-194"><label>[194]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>X</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>Champ: controllable and consistent human image animation with 3D parametric guidance</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>145</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-73001-6_9</pub-id>.</mixed-citation></ref>
<ref id="ref-195"><label>[195]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Novel view extrapolation with video diffusion priors</article-title>. <comment>arXiv:2411.14208. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.14208">http://arxiv.org/abs/2411.14208</ext-link>.</mixed-citation></ref>
<ref id="ref-196"><label>[196]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name></person-group>. <chapter-title>PhysGen: rigid-body physics-grounded image-to-video generation</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>360</fpage>&#x2013;<lpage>78</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-73007-8_21</pub-id>.</mixed-citation></ref>
<ref id="ref-197"><label>[197]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Dou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Si</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Diffusion as shader: 3D-aware video diffusion for versatile video generation control</article-title>. <comment>arXiv:2501.03847. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.03847">http://arxiv.org/abs/2501.03847</ext-link>.</mixed-citation></ref>
<ref id="ref-198"><label>[198]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>STAR: spatial-temporal augmentation with text-to-video models for real-world video super-resolution</article-title>. <comment>arXiv:2501.02976. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.02976">http://arxiv.org/abs/2501.02976</ext-link>.</mixed-citation></ref>
<ref id="ref-199"><label>[199]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Video depth anything: consistent depth estimation for super-long videos</article-title>. <comment>arXiv:2501.12375. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.12375">http://arxiv.org/abs/2501.12375</ext-link>.</mixed-citation></ref>
<ref id="ref-200"><label>[200]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yatim</surname> <given-names>D</given-names></string-name>, <string-name><surname>Fridman</surname> <given-names>R</given-names></string-name>, <string-name><surname>Bar-Tal</surname> <given-names>O</given-names></string-name>, <string-name><surname>Dekel</surname> <given-names>T</given-names></string-name></person-group>. <article-title>DynVFX: augmenting real videos with dynamic content</article-title>. <comment>arXiv:2502.03621. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.03621">http://arxiv.org/abs/2502.03621</ext-link>.</mixed-citation></ref>
<ref id="ref-201"><label>[201]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>W</given-names></string-name></person-group>. <article-title>FramePainter: endowing interactive image editing with video diffusion priors</article-title>. <comment>arXiv:2501.08225. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.08225">http://arxiv.org/abs/2501.08225</ext-link>.</mixed-citation></ref>
<ref id="ref-202"><label>[202]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Song</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DynamicFace: high-quality and consistent video face swapping using composable 3D facial priors</article-title>. <comment>arXiv:2501.08553. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.08553">http://arxiv.org/abs/2501.08553</ext-link>.</mixed-citation></ref>
<ref id="ref-203"><label>[203]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>CustomTTT: motion and appearance customized video generation via test-time training</article-title>. <comment>arXiv:2412.15646. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.15646">http://arxiv.org/abs/2412.15646</ext-link>.</mixed-citation></ref>
<ref id="ref-204"><label>[204]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pandey</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gadelha</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hold-Geoffroy</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>K</given-names></string-name>, <string-name><surname>Mitra</surname> <given-names>NJ</given-names></string-name>, <string-name><surname>Guerrero</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Motion modes: what could happen next?</article-title> <comment>arXiv:2412.00148. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.00148">http://arxiv.org/abs/2412.00148</ext-link>.</mixed-citation></ref>
<ref id="ref-205"><label>[205]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Diao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MoTrans: customized motion transfer with text-driven video diffusion models</article-title>. In: <conf-name>Proceedings of the 32nd ACM International Conference on Multimedia (MM &#x2019;24); 2024 Oct 28&#x2013;Nov 1</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2024</year>. p. <fpage>3421</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3664647.3680718</pub-id>.</mixed-citation></ref>
<ref id="ref-206"><label>[206]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yesiltepe</surname> <given-names>H</given-names></string-name>, <string-name><surname>Meral</surname> <given-names>THS</given-names></string-name>, <string-name><surname>Dunlop</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yanardag</surname> <given-names>P</given-names></string-name></person-group>. <article-title>MotionShop: zero-shot motion transfer in video diffusion models with mixture of score guidance</article-title>. <comment>arXiv:2412.05355. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.05355">http://arxiv.org/abs/2412.05355</ext-link>.</mixed-citation></ref>
<ref id="ref-207"><label>[207]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pondaven</surname> <given-names>A</given-names></string-name>, <string-name><surname>Siarohin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tulyakov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Torr</surname> <given-names>P</given-names></string-name>, <string-name><surname>Pizzati</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Video motion transfer with diffusion transformers</article-title>. <comment>arXiv:2412.07776. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.07776">http://arxiv.org/abs/2412.07776</ext-link>.</mixed-citation></ref>
<ref id="ref-208"><label>[208]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Loy</surname> <given-names>CC</given-names></string-name></person-group>. <article-title>ObjCtrl-2.5D: training-free object control with camera poses</article-title>. <comment>arXiv:2412.07721. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.07721">http://arxiv.org/abs/2412.07721</ext-link>.</mixed-citation></ref>
<ref id="ref-209"><label>[209]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>An</surname> <given-names>J</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Latent-Reframe: enabling camera control for video diffusion model without training</article-title>. <comment>arXiv:2412.06029. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.06029">http://arxiv.org/abs/2412.06029</ext-link>.</mixed-citation></ref>
<ref id="ref-210"><label>[210]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MotionBooth: motion-aware customized text-to-video generation</article-title>. <comment>arXiv:2406.17758. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2406.17758">http://arxiv.org/abs/2406.17758</ext-link>.</mixed-citation></ref>
<ref id="ref-211"><label>[211]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Reda</surname> <given-names>F</given-names></string-name>, <string-name><surname>Kontkanen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tabellion</surname> <given-names>E</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>D</given-names></string-name>, <string-name><surname>Pantofaru</surname> <given-names>C</given-names></string-name>, <string-name><surname>Curless</surname> <given-names>B</given-names></string-name></person-group>. <chapter-title>FILM: frame interpolation for large motion</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Avidan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Brostow</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ciss&#x00E9;</surname> <given-names>M</given-names></string-name>, <string-name><surname>Farinella</surname> <given-names>GM</given-names></string-name>, <string-name><surname>Hassner</surname> <given-names>T</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2022</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2022</year>. p. <fpage>250</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-20071-7_15</pub-id>.</mixed-citation></ref>
<ref id="ref-212"><label>[212]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jain</surname> <given-names>S</given-names></string-name>, <string-name><surname>Watson</surname> <given-names>D</given-names></string-name>, <string-name><surname>Tabellion</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ho?ynski</surname> <given-names>A</given-names></string-name>, <string-name><surname>Poole</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kontkanen</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Video interpolation with diffusion models</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>7341</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-213"><label>[213]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mahdizadehaghdam</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>W</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MaskINT: video editing via interpolative non-autoregressive masked transformers</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>7403</fpage>&#x2013;<lpage>12</lpage>.</mixed-citation></ref>
<ref id="ref-214"><label>[214]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ding</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Niklaus</surname> <given-names>S</given-names></string-name>, <string-name><surname>Abrevaya</surname> <given-names>V</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>Explorative inbetweening of time and space</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>378</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-73229-4_22</pub-id>.</mixed-citation></ref>
<ref id="ref-215"><label>[215]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xing</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>ToonCrafter: generative cartoon interpolation</article-title>. <source>ACM Trans Graph</source>. <year>2024</year>;<volume>43</volume>(<issue>6</issue>):<fpage>245:1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3687761</pub-id>.</mixed-citation></ref>
<ref id="ref-216"><label>[216]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bo</surname> <given-names>L</given-names></string-name></person-group>. <article-title>DiffuEraser: a diffusion model for video inpainting</article-title>. <comment>arXiv:2501.10018. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.10018">http://arxiv.org/abs/2501.10018</ext-link>.</mixed-citation></ref>
<ref id="ref-217"><label>[217]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SVFR: a unified framework for generalized video face restoration</article-title>. <comment>arXiv:2501.01235. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.01235">http://arxiv.org/abs/2501.01235</ext-link>.</mixed-citation></ref>
<ref id="ref-218"><label>[218]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Tao</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Loy</surname> <given-names>CC</given-names></string-name></person-group>. <article-title>MatAnyone: stable video matting with consistent memory propagation</article-title>. <comment>arXiv:2501.14677. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.14677">http://arxiv.org/abs/2501.14677</ext-link>.</mixed-citation></ref>
<ref id="ref-219"><label>[219]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SeedVR: seeding infinity in diffusion transformer towards generic video restoration</article-title>. <comment>arXiv:2501.01320. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.01320">http://arxiv.org/abs/2501.01320</ext-link>.</mixed-citation></ref>
<ref id="ref-220"><label>[220]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ju</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VideoPainter: any-length video inpainting and editing with plug-and-play context control</article-title>. <comment>arXiv:2503.05639. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.05639">http://arxiv.org/abs/2503.05639</ext-link>.</mixed-citation></ref>
<ref id="ref-221"><label>[221]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Teodoro</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gunawan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>SY</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>M</given-names></string-name></person-group>. <article-title>MIVE: new design and benchmark for multi-instance video editing</article-title>. <comment>arXiv:2412.12877. 2024 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.12877">http://arxiv.org/abs/2412.12877</ext-link>.</mixed-citation></ref>
<ref id="ref-222"><label>[222]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>K</given-names></string-name></person-group>. <article-title>3D convolutional neural networks for human action recognition</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2013</year>;<volume>35</volume>(<issue>1</issue>):<fpage>221</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2012.59</pub-id>; <pub-id pub-id-type="pmid">22392705</pub-id></mixed-citation></ref>
<ref id="ref-223"><label>[223]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>A</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Non-local neural networks</article-title>. <comment>arXiv:1711.07971. 2018 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1711.07971">http://arxiv.org/abs/1711.07971</ext-link>.</mixed-citation></ref>
<ref id="ref-224"><label>[224]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ning</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Video swin transformer</article-title>. <comment>arXiv:2106.13230. 2021 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2106.13230">http://arxiv.org/abs/2106.13230</ext-link>.</mixed-citation></ref>
<ref id="ref-225"><label>[225]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fei</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>TS</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Vitron: a unified pixel-level vision LLM for understanding, generating, segmenting, editing</article-title>. <comment>arXiv:2412.19806. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.19806">http://arxiv.org/abs/2412.19806</ext-link>.</mixed-citation></ref>
<ref id="ref-226"><label>[226]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Sa2VA: marrying SAM2 with LLaVA for dense grounded understanding of images and videos</article-title>. <comment>arXiv:2501.04001. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.04001">http://arxiv.org/abs/2501.04001</ext-link>.</mixed-citation></ref>
<ref id="ref-227"><label>[227]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <etal>et al.</etal></person-group> <chapter-title>VideoMamba: state space model for efficient video understanding</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Leonardis</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ricci</surname> <given-names>E</given-names></string-name>, <string-name><surname>Roth</surname> <given-names>S</given-names></string-name>, <string-name><surname>Russakovsky</surname> <given-names>O</given-names></string-name>, <string-name><surname>Sattler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>, editors. <source>Computer vision&#x2013;ECCV 2024</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>237</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-73347-5_14</pub-id>.</mixed-citation></ref>
<ref id="ref-228"><label>[228]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Divot: diffusion powers video tokenizer for comprehension and generation</article-title>. <comment>arXiv:2412.04432. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.04432">http://arxiv.org/abs/2412.04432</ext-link>.</mixed-citation></ref>
<ref id="ref-229"><label>[229]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <string-name><surname>van Steenkiste</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kurach</surname> <given-names>K</given-names></string-name>, <string-name><surname>Marinier</surname> <given-names>R</given-names></string-name>, <string-name><surname>Michalski</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gelly</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Towards accurate generative models of video: a new metric &#x0026; challenges</article-title>. <comment>arXiv:1812.01717. 2019 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1812.01717">http://arxiv.org/abs/1812.01717</ext-link>.</mixed-citation></ref>
<ref id="ref-230"><label>[230]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fischer</surname> <given-names>P</given-names></string-name>, <string-name><surname>Dosovitskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ilg</surname> <given-names>E</given-names></string-name>, <string-name><surname>H&#x00E4;usser</surname> <given-names>P</given-names></string-name>, <string-name><surname>Haz&#x0131;rba&#x015F;</surname> <given-names>C</given-names></string-name>, <string-name><surname>Golkov</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>FlowNet: learning optical flow with convolutional networks</article-title>. <comment>arXiv:1504.06852. 2015 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1504.06852">http://arxiv.org/abs/1504.06852</ext-link>.</mixed-citation></ref>
<ref id="ref-231"><label>[231]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lubin</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A visual discrimination model for imaging system design and evaluation. In: Vision models for target detection and recognition</article-title>. <comment>Series on Information Display; Volume 2</comment>, p. <fpage>245</fpage>&#x2013;<lpage>83</lpage>. <comment>World Scientific</comment>; <year>1995 [cited 2025 Apr 12]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://worldscientific.com/doi/abs/10.1142/9789812831200_0010#">https://worldscientific.com/doi/abs/10.1142/9789812831200_0010#</ext-link>.</mixed-citation></ref>
<ref id="ref-232"><label>[232]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hessel</surname> <given-names>J</given-names></string-name>, <string-name><surname>Holtzman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Forbes</surname> <given-names>M</given-names></string-name>, <string-name><surname>Bras</surname> <given-names>RL</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>CLIPScore: a reference-free evaluation metric for image captioning</article-title>. <comment>arXiv:2104.08718. 2022 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2104.08718">http://arxiv.org/abs/2104.08718</ext-link>.</mixed-citation></ref>
<ref id="ref-233"><label>[233]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Streijl</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Winkler</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hands</surname> <given-names>DS</given-names></string-name></person-group>. <article-title>Mean opinion score (MOS) revisited: methods and applications, limitations and alternatives</article-title>. <source>Multimed Syst</source>. <year>2016</year>;<volume>22</volume>(<issue>2</issue>):<fpage>213</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00530-014-0446-1</pub-id>.</mixed-citation></ref>
<ref id="ref-234"><label>[234]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Blattmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rombach</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ling</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dockhorn</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>SW</given-names></string-name>, <string-name><surname>Fidler</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Align your latents: high-resolution video synthesis with latent diffusion models</article-title>. <comment>arXiv:2304.08818. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2304.08818">http://arxiv.org/abs/2304.08818</ext-link>.</mixed-citation></ref>
<ref id="ref-235"><label>[235]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MotionCtrl: a unified and flexible motion controller for video generation</article-title>. In: <conf-name>ACM SIGGRAPH 2024 Conference Papers (SIGGRAPH &#x2019;24); 2024 Jul 27&#x2013;Aug 1</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>; <year>2024</year>. p. <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3641519.3657518</pub-id>.</mixed-citation></ref>
<ref id="ref-236"><label>[236]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aberman</surname> <given-names>K</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lischinski</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cohen-Or</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Unpaired motion style transfer from video to animation</article-title>. <source>ACM Trans Graph</source>. <year>2020</year>;<volume>39</volume>(<issue>4</issue>):<fpage>137</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3386569.3392469</pub-id>.</mixed-citation></ref>
<ref id="ref-237"><label>[237]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Guizilini</surname> <given-names>V</given-names></string-name>, <string-name><surname>Irshad</surname> <given-names>MZ</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Shakhnarovich</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ambrus</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Zero-shot novel view and depth synthesis with multi-view geometric diffusion</article-title>. <comment>arXiv:2501.18804. 2025 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.18804">http://arxiv.org/abs/2501.18804</ext-link>.</mixed-citation></ref>
<ref id="ref-238"><label>[238]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Efficient emotional adaptation for audio-driven talking-head generation</article-title>. <comment>arXiv:2309.04946. 2023 [cited 2025 Apr 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2309.04946">http://arxiv.org/abs/2309.04946</ext-link>.</mixed-citation></ref>
<ref id="ref-239"><label>[239]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google DeepMind. 2025</collab></person-group>. <article-title>Music AI Sandbox, now with new features and broader access</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://deepmind.google/discover/blog/music-ai-sandbox-now-with-new-features-and-broader-access/">https://deepmind.google/discover/blog/music-ai-sandbox-now-with-new-features-and-broader-access/</ext-link>.</mixed-citation></ref>
<ref id="ref-240"><label>[240]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Evans</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Parker</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Carr</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Zukowski</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Taylor</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pons</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Stable audio open</article-title>. <comment>arXiv:2407.14358. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2407.14358">http://arxiv.org/abs/2407.14358</ext-link>.</mixed-citation></ref>
<ref id="ref-241"><label>[241]</label><mixed-citation publication-type="other"><article-title>OpenAI Research | Milestone</article-title> <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://openai.com/research/index/milestone/">https://openai.com/research/index/milestone/</ext-link>.</mixed-citation></ref>
<ref id="ref-242"><label>[242]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Suno</surname> <given-names>AI</given-names></string-name></person-group>: <article-title>Suno AI V4 Is Hear [The Latest Version Unveiled]</article-title>. <comment>[cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://sunnoai.com/v4/">https://sunnoai.com/v4/</ext-link>.</mixed-citation></ref>
<ref id="ref-243"><label>[243]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google DeepMind. 2024</collab></person-group>. <article-title>Generating audio for video</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://deepmind.google/discover/blog/generating-audio-for-video/">https://deepmind.google/discover/blog/generating-audio-for-video/</ext-link>.</mixed-citation></ref>
<ref id="ref-244"><label>[244]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Copet</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kreuk</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gat</surname> <given-names>I</given-names></string-name>, <string-name><surname>Remez</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kant</surname> <given-names>D</given-names></string-name>, <string-name><surname>Synnaeve</surname> <given-names>G</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Simple and controllable music generation</article-title>. <comment>arXiv:2306.05284. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2306.05284">http://arxiv.org/abs/2306.05284</ext-link>.</mixed-citation></ref>
<ref id="ref-245"><label>[245]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ACE-Step: a step towards music generation foundation model</article-title>. <comment>arXiv:2506.00045. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2506.00045">http://arxiv.org/abs/2506.00045</ext-link>.</mixed-citation></ref>
<ref id="ref-246"><label>[246]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Park</surname> <given-names>DS</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Denk</surname> <given-names>TI</given-names></string-name>, <string-name><surname>Ly</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Noise2Music: text-conditioned music generation with diffusion models</article-title>. <comment>arXiv:2302.03917. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2302.03917">http://arxiv.org/abs/2302.03917</ext-link>.</mixed-citation></ref>
<ref id="ref-247"><label>[247]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ning</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>G</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DiffRhythm: blazingly fast and embarrassingly simple end-to-end full-length song generation with latent diffusion</article-title>. <comment>arXiv:2503.01183. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.01183">http://arxiv.org/abs/2503.01183</ext-link>.</mixed-citation></ref>
<ref id="ref-248"><label>[248]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hung</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Majumder</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Mehrish</surname> <given-names>A</given-names></string-name>, <string-name><surname>Valle</surname> <given-names>R</given-names></string-name>, <string-name><surname>Catanzaro</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>TangoFlux: super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization</article-title>. <comment>arXiv:2412.21037. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.21037">http://arxiv.org/abs/2412.21037</ext-link>.</mixed-citation></ref>
<ref id="ref-249"><label>[249]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Seed-Music: a unified framework for high quality and controlled music generation</article-title>. <comment>arXiv:2409.09214. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2409.09214">http://arxiv.org/abs/2409.09214</ext-link>.</mixed-citation></ref>
<ref id="ref-250"><label>[250]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>SL</given-names></string-name>, <string-name><surname>Donahue</surname> <given-names>C</given-names></string-name>, <string-name><surname>Watanabe</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bryan</surname> <given-names>NJ</given-names></string-name></person-group>. <article-title>Music controlNet: multiple time-varying controls for music generation</article-title>. <source>IEEE/ACM Transact Audio, Speech, Lang Process</source>. <year>2024</year>;<volume>32</volume>:<fpage>2692</fpage>&#x2013;<lpage>703</lpage>. doi:<pub-id pub-id-type="doi">10.1109/taslp.2024.3399026</pub-id>.</mixed-citation></ref>
<ref id="ref-251"><label>[251]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Suv&#x00E9;e</surname> <given-names>D</given-names></string-name>, <string-name><surname>Vanderperren</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jonckers</surname> <given-names>V</given-names></string-name></person-group>. <article-title>JAsCo: an aspect-oriented approach tailored for component based software development</article-title>. In: <conf-name>Proceedings of the 2nd International Conference on Aspect-Oriented Software Development (AOSD &#x2019;03); 2023 Mar 17&#x2013;21</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>21</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1145/643603.643606</pub-id>.</mixed-citation></ref>
<ref id="ref-252"><label>[252]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yuan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>H</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zang</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>YuE: scaling open foundation models for long-form music generation</article-title>. <comment>arXiv:2503.08638. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.08638">http://arxiv.org/abs/2503.08638</ext-link>.</mixed-citation></ref>
<ref id="ref-253"><label>[253]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Both ears wide open: towards language-driven spatial audio generation</article-title>. <comment>arXiv:2410.10676. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.10676">http://arxiv.org/abs/2410.10676</ext-link>.</mixed-citation></ref>
<ref id="ref-254"><label>[254]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ikemiya</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>G</given-names></string-name>, <string-name><surname>Murata</surname> <given-names>N</given-names></string-name>, <string-name><surname>Mart&#x00ED;nez-Ram&#x00ED;rez</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>WH</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MusicMagus: zero-shot text-to-music editing via diffusion models</article-title>. In: <conf-name>Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI &#x2019;24); 2024 Aug 3&#x2013;9</conf-name>; <publisher-loc>Jeju, Republic of Korea</publisher-loc>. p. <fpage>7805</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.24963/ijcai.2024/864</pub-id>.</mixed-citation></ref>
<ref id="ref-255"><label>[255]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wichern</surname> <given-names>G</given-names></string-name>, <string-name><surname>Germain</surname> <given-names>FG</given-names></string-name>, <string-name><surname>Khurana</surname> <given-names>S</given-names></string-name>, <string-name><surname>Roux</surname> <given-names>JL</given-names></string-name></person-group>. <article-title>SMITIN: self-monitored inference-time INtervention for generative music transformers</article-title>. <source>IEEE Open J Signal Process</source>. <year>2025</year>;<volume>6</volume>:<fpage>266</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ojsp.2025.3534686</pub-id>.</mixed-citation></ref>
<ref id="ref-256"><label>[256]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Haji-Ali</surname> <given-names>M</given-names></string-name>, <string-name><surname>Menapace</surname> <given-names>W</given-names></string-name>, <string-name><surname>Siarohin</surname> <given-names>A</given-names></string-name>, <string-name><surname>Skorokhodov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Canberk</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>KS</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AV-Link: temporally-aligned diffusion features for cross-modal audio-video generation</article-title>. <comment>arXiv:2412.15191. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.15191">http://arxiv.org/abs/2412.15191</ext-link>.</mixed-citation></ref>
<ref id="ref-257"><label>[257]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>HK</given-names></string-name>, <string-name><surname>Ishii</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hayakawa</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shibuya</surname> <given-names>T</given-names></string-name>, <string-name><surname>Schwing</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mitsufuji</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Taming multimodal joint training for high-quality video-to-audio synthesis</article-title>. <comment>arXiv:2412.15322. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.15322">http://arxiv.org/abs/2412.15322</ext-link>.</mixed-citation></ref>
<ref id="ref-258"><label>[258]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Seetharaman</surname> <given-names>P</given-names></string-name>, <string-name><surname>Russell</surname> <given-names>B</given-names></string-name>, <string-name><surname>Nieto</surname> <given-names>O</given-names></string-name>, <string-name><surname>Bourgin</surname> <given-names>D</given-names></string-name>, <string-name><surname>Owens</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Video-guided foley sound generation with multimodal controls</article-title>. <comment>arXiv:2411.17698. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.17698">http://arxiv.org/abs/2411.17698</ext-link>.</mixed-citation></ref>
<ref id="ref-259"><label>[259]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhuo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chengjing</surname> <given-names>W</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multimodal music generation with explicit bridges and retrieval augmentation</article-title>. <comment>arXiv:2412.09428. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.09428">http://arxiv.org/abs/2412.09428</ext-link>.</mixed-citation></ref>
<ref id="ref-260"><label>[260]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gramaccioni</surname> <given-names>RF</given-names></string-name>, <string-name><surname>Marinoni</surname> <given-names>C</given-names></string-name>, <string-name><surname>Postolache</surname> <given-names>E</given-names></string-name>, <string-name><surname>Comunit&#x00E0;</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cosmo</surname> <given-names>L</given-names></string-name>, <string-name><surname>Reiss</surname> <given-names>JD</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Stable-V2A: synthesis of synchronized sound effects with temporal and semantic controls</article-title>. <comment>arXiv:2412.15023. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.15023">http://arxiv.org/abs/2412.15023</ext-link>.</mixed-citation></ref>
<ref id="ref-261"><label>[261]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VidMusician: video-to-music generation with semantic-rhythmic alignment via hierarchical visual features</article-title>. <comment>arXiv:2412.15023. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.06296">http://arxiv.org/abs/2412.06296</ext-link>.</mixed-citation></ref>
<ref id="ref-262"><label>[262]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Song</surname> <given-names>G</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DREAM-Talk: diffusion-based realistic emotional audio-driven method for single image talking face generation</article-title>. <comment>arXiv:2312.13578. 2023 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2312.13578">http://arxiv.org/abs/2312.13578</ext-link>.</mixed-citation></ref>
<ref id="ref-263"><label>[263]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Song</surname> <given-names>R</given-names></string-name></person-group>. <article-title>LoVA: long-form video-to-audio generation</article-title>. In: <conf-name>ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2025 Apr 6&#x2013;11</conf-name>; <publisher-loc>Hyderabad, India</publisher-loc>: <publisher-name>ICASSP</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-264"><label>[264]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Comanducci</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bestagini</surname> <given-names>P</given-names></string-name>, <string-name><surname>Tubaro</surname> <given-names>S</given-names></string-name></person-group>. <article-title>FakeMusicCaps: a dataset for detection and attribution of synthetic music generated via text-to-music models</article-title>. <comment>arXiv:2409.10684. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2409.10684">http://arxiv.org/abs/2409.10684</ext-link>.</mixed-citation></ref>
<ref id="ref-265"><label>[265]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ciranni</surname> <given-names>R</given-names></string-name>, <string-name><surname>Mariani</surname> <given-names>G</given-names></string-name>, <string-name><surname>Mancusi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Postolache</surname> <given-names>E</given-names></string-name>, <string-name><surname>Fabbro</surname> <given-names>G</given-names></string-name>, <string-name><surname>Rodol&#x00E0;</surname> <given-names>E</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>COCOLA: coherence-oriented contrastive learning of musical audio representations</article-title>. In: <conf-name>ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2025 Apr 6&#x2013;11</conf-name>; <publisher-loc>Hyderabad, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-266"><label>[266]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ni</surname> <given-names>J</given-names></string-name>, <string-name><surname>Song</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ghosal</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Yue</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MixEval-X: any-to-any evaluations from real-world data mixtures</article-title>. <comment>arXiv:2410.13754. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.13754">http://arxiv.org/abs/2410.13754</ext-link>.</mixed-citation></ref>
<ref id="ref-267"><label>[267]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A survey on deep learning for symbolic music generation: representations, algorithms, evaluations, and challenges</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>56</volume>(<issue>1</issue>):<fpage>7:1</fpage>&#x2013;<lpage>39</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3597493</pub-id>.</mixed-citation></ref>
<ref id="ref-268"><label>[268]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Lerch</surname> <given-names>A</given-names></string-name></person-group>. <article-title>On the evaluation of generative models in music</article-title>. <source>Neural Comput &#x0026; Applic</source>. <year>2020</year>;<volume>32</volume>(<issue>9</issue>):<fpage>4773</fpage>&#x2013;<lpage>84</lpage>.</mixed-citation></ref>
<ref id="ref-269"><label>[269]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>New evaluation methods for automatic music generation [Ph.D thesis]. University of York</article-title>; <year>2022</year> <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://etheses.whiterose.ac.uk/id/eprint/31507/">https://etheses.whiterose.ac.uk/id/eprint/31507/</ext-link>.</mixed-citation></ref>
<ref id="ref-270"><label>[270]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>CNN based music emotion classification</article-title>. <comment>arXiv:1704.05665. 2017 [cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1704.05665">http://arxiv.org/abs/1704.05665</ext-link>.</mixed-citation></ref>
<ref id="ref-271"><label>[271]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yading</surname> <given-names>S</given-names></string-name>, <string-name><surname>Simon</surname> <given-names>D</given-names></string-name>, <string-name><surname>Marcus</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Evaluation of musical features for emotion classification</article-title>. In: <conf-name>13th International Society for Music Information Retrieval Conference; 2012 Oct 8&#x2013;12</conf-name>; <publisher-loc>Porto, Portugal</publisher-loc>. p. <fpage>523</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-272"><label>[272]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hawthorne</surname> <given-names>C</given-names></string-name>, <string-name><surname>Simon</surname> <given-names>I</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zeghidour</surname> <given-names>N</given-names></string-name>, <string-name><surname>Gardner</surname> <given-names>J</given-names></string-name>, <string-name><surname>Manilow</surname> <given-names>E</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Multi-instrument music synthesis with spectrogram diffusion</article-title>. <comment>arXiv:2206.05408. 2022 [cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2206.05408">http://arxiv.org/abs/2206.05408</ext-link>.</mixed-citation></ref>
<ref id="ref-273"><label>[273]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Application of audio signal processing technology in music synthesis</article-title>. <source>J Elect Syst</source>. <year>2024</year>;<volume>20</volume>(<issue>9s</issue>):<fpage>648</fpage>&#x2013;<lpage>54</lpage>.</mixed-citation></ref>
<ref id="ref-274"><label>[274]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Valevski</surname> <given-names>D</given-names></string-name>, <string-name><surname>Leviathan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Arar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fruchter</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Diffusion models are real-time game engines</article-title>. <comment>arXiv:2408.14837. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2408.14837">http://arxiv.org/abs/2408.14837</ext-link>.</mixed-citation></ref>
<ref id="ref-275"><label>[275]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Oasis: A Universe in a Transformer</collab></person-group>. <article-title>Oasis: a universe in a transformer</article-title>. <comment>[cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.decart.ai/articles/oasis-interactive-ai-video-game-model">https://www.decart.ai/articles/oasis-interactive-ai-video-game-model</ext-link>.</mixed-citation></ref>
<ref id="ref-276"><label>[276]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Playable game generation</article-title>. <comment>arXiv:2412.00887. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.00887">http://arxiv.org/abs/2412.00887</ext-link>.</mixed-citation></ref>
<ref id="ref-277"><label>[277]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google DeepMind</collab></person-group>. <article-title>2025. Genie 2: A large-scale foundation world model</article-title>. <comment>[cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/">https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/</ext-link>.</mixed-citation></ref>
<ref id="ref-278"><label>[278]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Goel</surname> <given-names>V</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>G</given-names></string-name>, <string-name><surname>Korolev</surname> <given-names>S</given-names></string-name>, <string-name><surname>Terzopoulos</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Wonderland: navigating 3D scenes from a single image</article-title>. <comment>arXiv:2412.12091. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.12091">http://arxiv.org/abs/2412.12091</ext-link>.</mixed-citation></ref>
<ref id="ref-279"><label>[279]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DimensionX: create any 3D and 4D scenes from a single image with controllable video diffusion</article-title>. <comment>arXiv:2411.04928. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.04928">http://arxiv.org/abs/2411.04928</ext-link>.</mixed-citation></ref>
<ref id="ref-280"><label>[280]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tucker</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guibas</surname> <given-names>L</given-names></string-name>, <string-name><surname>Snavely</surname> <given-names>N</given-names></string-name>, <string-name><surname>Wetzstein</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Streetscapes: large-scale consistent street view generation using autoregressive video diffusion</article-title>. In: <conf-name>ACM SIGGRAPH 2024 Conference Papers; 2024 Jul 27&#x2013;Aug 1</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>11</lpage>. (SIGGRAPH &#x2019;24) doi:<pub-id pub-id-type="doi">10.1145/3641519.3657513</pub-id>.</mixed-citation></ref>
<ref id="ref-281"><label>[281]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Gupta</surname> <given-names>V</given-names></string-name>, <string-name><surname>Man</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>YX</given-names></string-name></person-group>. <article-title>PaintScene4D: consistent 4D scene generation from text prompts</article-title>. <comment>arXiv:2412.04471. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.04471">http://arxiv.org/abs/2412.04471</ext-link>.</mixed-citation></ref>
<ref id="ref-282"><label>[282]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SynCamMaster: synchronizing multi-camera video generation from diverse viewpoints</article-title>. <comment>arXiv:2412.07760. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.07760">http://arxiv.org/abs/2412.07760</ext-link>.</mixed-citation></ref>
<ref id="ref-283"><label>[283]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Kiss3DGen: repurposing image diffusion models for 3D asset generation</article-title>. <comment>arXiv:2503.01370. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.01370">http://arxiv.org/abs/2503.01370</ext-link>.</mixed-citation></ref>
<ref id="ref-284"><label>[284]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ling</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Light-a-video: training-free video relighting via progressive light fusion</article-title>. <comment>arXiv:2502.08590. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.08590">http://arxiv.org/abs/2502.08590</ext-link>.</mixed-citation></ref>
<ref id="ref-285"><label>[285]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>YJ</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Ke</surname> <given-names>BH</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>CWT</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AuraFusion360: augmented unseen region alignment for reference-based 360&#x00B0; unbounded scene inpainting</article-title>. <comment>arXiv:2502.05176. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.05176">http://arxiv.org/abs/2502.05176</ext-link>.</mixed-citation></ref>
<ref id="ref-286"><label>[286]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Condor</surname> <given-names>J</given-names></string-name>, <string-name><surname>Speierer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bode</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bozic</surname> <given-names>A</given-names></string-name>, <string-name><surname>Green</surname> <given-names>S</given-names></string-name>, <string-name><surname>Didyk</surname> <given-names>P</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Don&#x2019;t splat your gaussians: volumetric ray-traced primitives for modeling and rendering scattering and emissive media</article-title>. <source>ACM Trans Graph</source>. <year>2025</year>;<volume>44</volume>(<issue>1</issue>):<fpage>10:1</fpage>&#x2013;<lpage>10:17</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3711853</pub-id>.</mixed-citation></ref>
<ref id="ref-287"><label>[287]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sax</surname> <given-names>A</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>KJ</given-names></string-name>, <string-name><surname>Henaff</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fast3R: towards 3D reconstruction of 1000&#x002B; images in one forward pass</article-title>. <comment>arXiv:2501.13928. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.13928">http://arxiv.org/abs/2501.13928</ext-link>.</mixed-citation></ref>
<ref id="ref-288"><label>[288]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Stearns</surname> <given-names>C</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VideoLifter: lifting videos to 3D with fast hierarchical stereo alignment</article-title>. <comment>arXiv:2501.01949. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.01949">http://arxiv.org/abs/2501.01949</ext-link>.</mixed-citation></ref>
<ref id="ref-289"><label>[289]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lyu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kundu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tsai</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>MH</given-names></string-name></person-group>. <article-title>Gaga: group any gaussians via 3D-aware memory bank</article-title>. <comment>arXiv:2404.07977. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2404.07977">http://arxiv.org/abs/2404.07977</ext-link>.</mixed-citation></ref>
<ref id="ref-290"><label>[290]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Qiu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AniGS: animatable gaussian avatar from a single image with inconsistent gaussian reconstruction</article-title>. <comment>arXiv:2412.02684. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.02684">http://arxiv.org/abs/2412.02684</ext-link>.</mixed-citation></ref>
<ref id="ref-291"><label>[291]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kant</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Weber</surname> <given-names>E</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Khirodkar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhaoen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Martinez</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Pippo: high-resolution multi-view humans from a single image</article-title>. <comment>arXiv:2502.07785. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2502.07785">http://arxiv.org/abs/2502.07785</ext-link>.</mixed-citation></ref>
<ref id="ref-292"><label>[292]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cha</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>I</given-names></string-name>, <string-name><surname>Joo</surname> <given-names>H</given-names></string-name></person-group>. <article-title>PERSE: personalized 3D generative avatars from a single portrait</article-title>. <comment>arXiv:2412.21206. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.21206">http://arxiv.org/abs/2412.21206</ext-link>.</mixed-citation></ref>
<ref id="ref-293"><label>[293]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sundaresan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Sadigh</surname> <given-names>D</given-names></string-name>, <string-name><surname>Choudhury</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bohg</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Motion tracks: a unified representation for human-robot transfer in few-shot imitation learning</article-title>. <comment>arXiv:2501.06994. 2025 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.06994">http://arxiv.org/abs/2501.06994</ext-link>.</mixed-citation></ref>
<ref id="ref-294"><label>[294]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>JS</given-names></string-name>, <string-name><surname>O&#x2019;Brien</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Morris</surname> <given-names>MR</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bernstein</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Generative agents: interactive simulacra of human behavior</article-title>. In: <conf-name>Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST &#x2019;23); 2023 Oct 29&#x2013;Nov 1</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3586183.3606763</pub-id>.</mixed-citation></ref>
<ref id="ref-295"><label>[295]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google DeepMind</collab></person-group>. <article-title>2025. A generalist AI agent for 3D virtual environments</article-title>. <comment>[cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://deepmind.google/discover/blog/sima-generalist-ai-agent-for-3d-virtual-environments/">https://deepmind.google/discover/blog/sima-generalist-ai-agent-for-3d-virtual-environments/</ext-link>.</mixed-citation></ref>
<ref id="ref-296"><label>[296]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Dynamic try-on: taming video virtual try-on with dynamic attention mechanism</article-title>. <comment>arXiv:2412.09822. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.09822">http://arxiv.org/abs/2412.09822</ext-link>.</mixed-citation></ref>
<ref id="ref-297"><label>[297]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Du</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Stag-1: towards realistic 4D driving simulation with video generation model</article-title>. <comment>arXiv:2412.05280. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.05280">http://arxiv.org/abs/2412.05280</ext-link>.</mixed-citation></ref>
<ref id="ref-298"><label>[298]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>GiiNEX</collab></person-group>. <article-title>[cited 2025 Apr 13]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://giinex.tencent.com/#/index">https://giinex.tencent.com/#/index</ext-link>.</mixed-citation></ref>
<ref id="ref-299"><label>[299]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kanervisto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bignell</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wen</surname> <given-names>LY</given-names></string-name>, <string-name><surname>Grayson</surname> <given-names>M</given-names></string-name>, <string-name><surname>Georgescu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Valcarcel Macua</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>World and human action models towards gameplay ideation</article-title>. <source>Nature</source>. <year>2025</year>;<volume>638</volume>(<issue>8051</issue>):<fpage>656</fpage>&#x2013;<lpage>63</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41586-025-08600-3</pub-id>; <pub-id pub-id-type="pmid">39972228</pub-id></mixed-citation></ref>
<ref id="ref-300"><label>[300]</label><mixed-citation publication-type="other"><article-title>Layer | Game Art Without Limits</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://layer.ai/">https://layer.ai/</ext-link>.</mixed-citation></ref>
<ref id="ref-301"><label>[301]</label><mixed-citation publication-type="other"><article-title>Scenario - Take complete control of your AI workflows</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.scenario.com/">https://www.scenario.com/</ext-link>.</mixed-citation></ref>
<ref id="ref-302"><label>[302]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Gemini</collab></person-group>. <article-title>Gemini Apps&#x2019; release updates &#x0026; improvements</article-title>. <comment>[cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://gemini.google/release-notes/">https://gemini.google/release-notes/</ext-link>.</mixed-citation></ref>
<ref id="ref-303"><label>[303]</label><mixed-citation publication-type="other"><article-title>Claude 3.7 Sonnet and Claude Code</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.anthropic.com/news/claude-3-7-sonnet">https://www.anthropic.com/news/claude-3-7-sonnet</ext-link>.</mixed-citation></ref>
<ref id="ref-304"><label>[304]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>DeepSeek-AI</collab>, <string-name><surname>Guo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>DeepSeek-R1: incentivizing reasoning capability in LLMs via reinforcement learning</article-title>. <comment>arXiv:2501.12948. 2025 [cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.12948">http://arxiv.org/abs/2501.12948</ext-link>.</mixed-citation></ref>
<ref id="ref-305"><label>[305]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>mistralai/Mistral-Large-Instruct-2407</collab></person-group>. <article-title>Hugging Face</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://huggingface.co/mistralai/Mistral-Large-Instruct-2407">https://huggingface.co/mistralai/Mistral-Large-Instruct-2407</ext-link>.</mixed-citation></ref>
<ref id="ref-306"><label>[306]</label><mixed-citation publication-type="other"><article-title>gpt-4-5-system-card-2272025.pdf</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf">https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-307"><label>[307]</label><mixed-citation publication-type="other"><article-title>News | xAI</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://x.ai/news">https://x.ai/news</ext-link>.</mixed-citation></ref>
<ref id="ref-308"><label>[308]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Agrawal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Antoniak</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hanna</surname> <given-names>EB</given-names></string-name>, <string-name><surname>Bout</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chaplot</surname> <given-names>D</given-names></string-name>, <string-name><surname>Chudnovsky</surname> <given-names>J</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Pixtral 12B</article-title>. <comment>arXiv:2410.07073. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2410.07073">http://arxiv.org/abs/2410.07073</ext-link>.</mixed-citation></ref>
<ref id="ref-309"><label>[309]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>meta-llama/llama3</collab></person-group>. <article-title>Meta Llama</article-title>; <year>2025</year> <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/meta-llama/llama3">https://github.com/meta-llama/llama3</ext-link>.</mixed-citation></ref>
<ref id="ref-310"><label>[310]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>AQ</given-names></string-name>, <string-name><surname>Sablayrolles</surname> <given-names>A</given-names></string-name>, <string-name><surname>Roux</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mensch</surname> <given-names>A</given-names></string-name>, <string-name><surname>Savary</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bamford</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Mixtral of experts</article-title>. <comment>arXiv:2401.04088. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2401.04088">http://arxiv.org/abs/2401.04088</ext-link>.</mixed-citation></ref>
<ref id="ref-311"><label>[311]</label><mixed-citation publication-type="other"><article-title>License and copyright - arXiv info</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://info.arxiv.org/help/license/index.html#licenses-available">https://info.arxiv.org/help/license/index.html#licenses-available</ext-link>.</mixed-citation></ref>
<ref id="ref-312"><label>[312]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Team</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kamath</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ferret</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pathak</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vieillard</surname> <given-names>N</given-names></string-name>, <string-name><surname>Merhej</surname> <given-names>R</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Gemma 3 technical report</article-title>. <comment>arXiv:2503.19786. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2503.19786">http://arxiv.org/abs/2503.19786</ext-link>.</mixed-citation></ref>
<ref id="ref-313"><label>[313]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Abdin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Aneja</surname> <given-names>J</given-names></string-name>, <string-name><surname>Behl</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bubeck</surname> <given-names>S</given-names></string-name>, <string-name><surname>Eldan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Gunasekar</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Phi-4 technical report</article-title>. <comment>arXiv:2412.08905. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.08905">http://arxiv.org/abs/2412.08905</ext-link>.</mixed-citation></ref>
<ref id="ref-314"><label>[314]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Databricks</collab></person-group>. <article-title>2024 introducing DBRX: a new state-of-the-art open LLM</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm">https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm</ext-link>.</mixed-citation></ref>
<ref id="ref-315"><label>[315]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Nvidia</collab>, <string-name><surname>Adler</surname> <given-names>B</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>N</given-names></string-name>, <string-name><surname>Aithal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Anh</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Bhattacharya</surname> <given-names>P</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Nemotron-4 340B technical report</article-title>. <comment>arXiv:2406.11704. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2406.11704">http://arxiv.org/abs/2406.11704</ext-link>.</mixed-citation></ref>
<ref id="ref-316"><label>[316]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Team</surname> <given-names>TH</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Hunyuan-TurboS: advancing large language models through mamba-transformer synergy and adaptive chain-of-thought</article-title>. <comment>arXiv:2505.15431. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2505.15431">http://arxiv.org/abs/2505.15431</ext-link>.</mixed-citation></ref>
<ref id="ref-317"><label>[317]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>llm.hunyuan.T1</collab></person-group>. <article-title>llm.hunyuan.T1</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://tencent.github.io/llm.hunyuan.T1/README_EN.html">https://tencent.github.io/llm.hunyuan.T1/README_EN.html</ext-link>.</mixed-citation></ref>
<ref id="ref-318"><label>[318]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>GitHub</collab></person-group>. <article-title>Build software better, together</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com">https://github.com</ext-link>.</mixed-citation></ref>
<ref id="ref-319"><label>[319]</label><mixed-citation publication-type="other"><article-title>Kimi K2: open agentic intelligence</article-title>. <comment>[cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://moonshotai.github.io/Kimi-K2/">https://moonshotai.github.io/Kimi-K2/</ext-link>.</mixed-citation></ref>
<ref id="ref-320"><label>[320]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>AlibabaCloud</collab></person-group>. <article-title>Tongyi Qianwen (Qwen) - Alibaba Cloud</article-title>. <comment>[cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.alibabacloud.com/solutions/generative-ai/qwen">https://www.alibabacloud.com/solutions/generative-ai/qwen</ext-link>.</mixed-citation></ref>
<ref id="ref-321"><label>[321]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hui</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Qwen2 technical report</article-title>. <comment>arXiv:2407.10671. 2024 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2407.10671">http://arxiv.org/abs/2407.10671</ext-link>.</mixed-citation></ref>
<ref id="ref-322"><label>[322]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>MiniMax</collab>, <string-name><surname>Li</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>MiniMax-01: scaling foundation models with lightning attention</article-title>. <comment>arXiv:2501.08313. 2025 [cited 2025 Jul 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2501.08313">http://arxiv.org/abs/2501.08313</ext-link>.</mixed-citation></ref>
<ref id="ref-323"><label>[323]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Song</surname> <given-names>S</given-names></string-name>, <string-name><surname>Duah</surname> <given-names>B</given-names></string-name>, <string-name><surname>Macbeth</surname> <given-names>J</given-names></string-name>, <string-name><surname>Carter</surname> <given-names>S</given-names></string-name>, <string-name><surname>Van</surname> <given-names>MP</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>More human than human: LLM-generated narratives outperform human-LLM interleaved narratives</article-title>. In: <conf-name>Proceedings of the 15th Conference on Creativity and Cognition; 2023 Jun 19&#x2013;21</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>368</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3591196.3596612</pub-id>.</mixed-citation></ref>
<ref id="ref-324"><label>[324]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Google</collab></person-group>. <article-title>2024. Introducing Gemini 2.0: our new AI model for the agentic era</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/">https://blog.google/technology/google-deepmind/google-gemini-ai-update-december-2024/</ext-link>.</mixed-citation></ref>
<ref id="ref-325"><label>[325]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Anand</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Unveiling google&#x2019;s gemini 2.0: a comprehensive study of its multimodal AI design, advanced architecture, and real-world applications</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/publication/387089907_Unveiling_Google&#x2019;s_Gemini_20_A_Comprehensive_Study_of_its_Multimodal_AI_Design_Advanced_Architecture_and_Real-World_Applications">https://www.researchgate.net/publication/387089907_Unveiling_Google&#x2019;s_Gemini_20_A_Comprehensive_Study_of_its_Multimodal_AI_Design_Advanced_Architecture_and_Real-World_Applications</ext-link>.</mixed-citation></ref>
<ref id="ref-326"><label>[326]</label><mixed-citation publication-type="other"><article-title>Introducing Claude 3.5 Sonnet</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.anthropic.com/news/claude-3-5-sonnet">https://www.anthropic.com/news/claude-3-5-sonnet</ext-link>.</mixed-citation></ref>
<ref id="ref-327"><label>[327]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Virtual YouTuber Wiki</collab></person-group>. <article-title>2025. Neuro-sama</article-title>. <comment>[cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://virtualyoutuber.fandom.com/wiki/Neuro-sama">https://virtualyoutuber.fandom.com/wiki/Neuro-sama</ext-link>.</mixed-citation></ref>
<ref id="ref-328"><label>[328]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Magic-Emerge/luna-ai</collab></person-group>. <article-title>Magic Emerge</article-title>. <comment>2025 [cited 2025 Apr 13]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/Magic-Emerge/luna-ai">https://github.com/Magic-Emerge/luna-ai</ext-link>.</mixed-citation></ref>
<ref id="ref-329"><label>[329]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Desai</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Ravi</surname> <given-names>T</given-names></string-name>, <string-name><surname>Luqman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kota</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yadav</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Gen-AI for user safety: a survey</article-title>. In: <conf-name>2024 IEEE International Conference on Big Data (BigData); 2024 Dec 15&#x2013;18</conf-name>; <publisher-loc>Washington, DC, USA</publisher-loc>. p. <fpage>5315</fpage>&#x2013;<lpage>24</lpage>.</mixed-citation></ref>
<ref id="ref-330"><label>[330]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Karagoz</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Ethics and technical aspects of generative AI models in digital content creation</article-title>. <comment>arXiv:2412.16389. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.16389">http://arxiv.org/abs/2412.16389</ext-link>.</mixed-citation></ref>
<ref id="ref-331"><label>[331]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Villate-Castillo</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ser</surname> <given-names>JD</given-names></string-name>, <string-name><surname>Sanz</surname> <given-names>B</given-names></string-name></person-group>. <article-title>A collaborative content moderation framework for toxicity detection based on conformalized estimates of annotation disagreement</article-title>. <comment>arXiv:2411.04090. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2411.04090">http://arxiv.org/abs/2411.04090</ext-link>.</mixed-citation></ref>
<ref id="ref-332"><label>[332]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Banchio</surname> <given-names>PR</given-names></string-name></person-group>. <article-title>Legal, ethical and practical challenges of AI-driven content moderation. Rochester, NY: Social Science Research Network</article-title>. <comment>2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://papers.ssrn.com/abstract=4984756">https://papers.ssrn.com/abstract=4984756</ext-link>.</mixed-citation></ref>
<ref id="ref-333"><label>[333]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lloyd</surname> <given-names>T</given-names></string-name>, <string-name><surname>Reagle</surname> <given-names>J</given-names></string-name>, <string-name><surname>Naaman</surname> <given-names>M</given-names></string-name></person-group>. <article-title>&#x201C;There has to be a lot that we&#x2019;re missing&#x201D;: moderating AI-generated content on reddit</article-title>. <comment>arXiv:2311.12702. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2311.12702">http://arxiv.org/abs/2311.12702</ext-link>.</mixed-citation></ref>
<ref id="ref-334"><label>[334]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Abdali</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anarfi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Barberan</surname> <given-names>C</given-names></string-name>, <string-name><surname>He</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Decoding the AI pen: techniques and challenges in detecting AI-generated text</article-title>. In: <conf-name>Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD &#x2019;24); 2024 Aug 25&#x2013;29</conf-name>; <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>. p. <fpage>6428</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3637528.3671463</pub-id>.</mixed-citation></ref>
<ref id="ref-335"><label>[335]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Han</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Synthetic data in AI: challenges, applications, and ethical implications</article-title>. <comment>arXiv:2401.01629. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2401.01629">http://arxiv.org/abs/2401.01629</ext-link>.</mixed-citation></ref>
<ref id="ref-336"><label>[336]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Appel</surname> <given-names>RE</given-names></string-name></person-group>. <article-title>Generative AI regulation can learn from social media regulation</article-title>. <comment>arXiv:2412.11335. 2024 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2412.11335">http://arxiv.org/abs/2412.11335</ext-link>.</mixed-citation></ref>
<ref id="ref-337"><label>[337]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hagerty</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rubinov</surname> <given-names>I</given-names></string-name></person-group>. <article-title>Global AI ethics: a review of the social impacts and ethical implications of artificial intelligence</article-title>. <comment>arXiv:1907.07892. 2019 [cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1907.07892">http://arxiv.org/abs/1907.07892</ext-link>.</mixed-citation></ref>
<ref id="ref-338"><label>[338]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Arora</surname> <given-names>S</given-names></string-name>, <string-name><surname>Arora</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hastings</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The psychological impacts of algorithmic and AI-driven social media on teenagers: a call to action</article-title>. In: <conf-name>2024 IEEE Digital Platforms and Societal Harms (DPSH); 2024 Oct 14&#x2013;15</conf-name>; <publisher-loc>Washington, DC, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/dpsh60098.2024.10774922</pub-id>.</mixed-citation></ref>
<ref id="ref-339"><label>[339]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Khlaif</surname> <given-names>ZN</given-names></string-name></person-group>. <article-title>Ethical concerns about using AI-generated text in scientific research. Rochester, NY: Social Science Research Network</article-title>; <year>2023</year> <comment>[cited 2025 Apr 8]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://papers.ssrn.com/abstract=4387984">https://papers.ssrn.com/abstract=4387984</ext-link>.</mixed-citation></ref>
<ref id="ref-340"><label>[340]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Illia</surname> <given-names>L</given-names></string-name>, <string-name><surname>Colleoni</surname> <given-names>E</given-names></string-name>, <string-name><surname>Zyglidopoulos</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Ethical implications of text generation in the age of artificial intelligence</article-title>. <source>Business Ethics Environ Respons</source>. <year>2023</year>;<volume>32</volume>(<issue>1</issue>):<fpage>201</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-341"><label>[341]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gillespie</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Content moderation, AI, and the question of scale</article-title>. <source>Big Data &#x0026; Society [Internet]</source>. <year>2020 Aug 21 [cited 2025 Apr 8]</year>;<volume>7</volume>(<issue>2</issue>). Available from: <ext-link ext-link-type="uri" xlink:href="https://journals.sagepub.com/doi/full/10.1177/2053951720943234">https://journals.sagepub.com/doi/full/10.1177/2053951720943234</ext-link>.</mixed-citation></ref>
<ref id="ref-342"><label>[342]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Stout</surname> <given-names>DW</given-names></string-name></person-group>. <article-title>How generative AI Has transformed creative work: a comprehensive study. Magai</article-title>. <year>2025</year> <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://magai.co/generative-ai-has-transformed-creative-work/">https://magai.co/generative-ai-has-transformed-creative-work/</ext-link>.</mixed-citation></ref>
<ref id="ref-343"><label>[343]</label><mixed-citation publication-type="other"><article-title>50 arguments against the use of AI in creative fields</article-title>. <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://aokistudio.com/50-arguments-against-the-use-of-ai-in-creative-fields.html">https://aokistudio.com/50-arguments-against-the-use-of-ai-in-creative-fields.html</ext-link>.</mixed-citation></ref>
<ref id="ref-344"><label>[344]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thambaiya</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kariyawasam</surname> <given-names>K</given-names></string-name>, <string-name><surname>Talagala</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Copyright law in the age of AI: analysing the AI-generated works and copyright challenges in Australia</article-title>. <source>Int Rev Law Comput Technol</source>. <year>2025</year>;<volume>35</volume>(<issue>2</issue>):<fpage>1</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1080/13600869.2025.2486893</pub-id>.</mixed-citation></ref>
<ref id="ref-345"><label>[345]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Topics | European Parliament. 2023</collab></person-group>. <article-title>EU AI Act: first regulation on artificial intelligence</article-title>. <comment>[cited 2025 Jun 4]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence">https://www.europarl.europa.eu/topics/en/article/20230601STO93804/eu-ai-act-first-regulation-on-artificial-intelligence</ext-link>.</mixed-citation></ref>
<ref id="ref-346"><label>[346]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>GOV.UK</collab></person-group>. <article-title>Artificial Intelligence and Intellectual Property: copyright and patents</article-title>. <comment>[cited 2025 Jun 4]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.gov.uk/government/consultations/artificial-intelligence-and-ip-copyright-and-patents/artificial-intelligence-and-intellectual-property-copyright-and-patents">https://www.gov.uk/government/consultations/artificial-intelligence-and-ip-copyright-and-patents/artificial-intelligence-and-intellectual-property-copyright-and-patents</ext-link>.</mixed-citation></ref>
<ref id="ref-347"><label>[347]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>kimjones</collab></person-group>. <article-title>Adapting models to handle cultural variations in language and context. Welocalize</article-title>. <year>2024</year> <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.welocalize.com/insights/adapting-models-to-handle-cultural-variations-in-language-and-context/">https://www.welocalize.com/insights/adapting-models-to-handle-cultural-variations-in-language-and-context/</ext-link>.</mixed-citation></ref>
<ref id="ref-348"><label>[348]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Trinity College Dublin</collab></person-group>. <article-title>Generative AI models are encoding biases and negative stereotypes in their users</article-title>. <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.tcd.ie/news_events/articles/2023/generative-ai-models-are-encoding-biases-and-negative-stereotypes-in-their-users/">https://www.tcd.ie/news_events/articles/2023/generative-ai-models-are-encoding-biases-and-negative-stereotypes-in-their-users/</ext-link>.</mixed-citation></ref>
<ref id="ref-349"><label>[349]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Spotify</collab></person-group>. <article-title>Long Reads</article-title>. <comment>[cited 2025 Jun 4]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://open.spotify.com/show/4mqFZZV8DMVeN80VwUumP4">https://open.spotify.com/show/4mqFZZV8DMVeN80VwUumP4</ext-link>.</mixed-citation></ref>
<ref id="ref-350"><label>[350]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Studio Ghibli AI Art Sparks Big Debate Over Creativity and Copyright</collab></person-group>. <article-title>Studio ghibli AI art sparks big debate over creativity and copyright</article-title>. <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://opentools.ai/news/studio-ghibli-ai-art-sparks-big-debate-over-creativity-and-copyright">https://opentools.ai/news/studio-ghibli-ai-art-sparks-big-debate-over-creativity-and-copyright</ext-link>.</mixed-citation></ref>
<ref id="ref-351"><label>[351]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Castro</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Critics of generative AI are worrying about the wrong IP issues</article-title>. <year>2023 Mar [cited 2025 May 20]</year>. Available from: <ext-link ext-link-type="uri" xlink:href="https://itif.org/publications/2023/03/20/critics-of-generative-ai-are-worrying-about-the-wrong-ip-issues/">https://itif.org/publications/2023/03/20/critics-of-generative-ai-are-worrying-about-the-wrong-ip-issues/</ext-link>.</mixed-citation></ref>
<ref id="ref-352"><label>[352]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Rice</surname> <given-names>AB</given-names></string-name></person-group>. <article-title>Miller library libguides: AI as a research tool: ethical and legal considerations</article-title>. <comment>[cited 2025 Jun 4]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://washcoll.libguides.com/c.php?g=1414339&#x0026;p=10478589">https://washcoll.libguides.com/c.php?g=1414339&#x0026;p=10478589</ext-link>.</mixed-citation></ref>
<ref id="ref-353"><label>[353]</label><mixed-citation publication-type="other"><article-title>How animation industry can be transformed by generative AI</article-title>. <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.e2enetworks.com/blog/how-animation-industry-can-be-transformed-by-generative-ai">https://www.e2enetworks.com/blog/how-animation-industry-can-be-transformed-by-generative-ai</ext-link>.</mixed-citation></ref>
<ref id="ref-354"><label>[354]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>AI Takes Center Stage in Animation&#x2019;s Future: Top Tools of 2024 Revealed!</collab></person-group>. <article-title>AI Takes Center Stage in Animation&#x2019;s Future: Top Tools of 2024 Revealed!</article-title>. <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://opentools.ai/news/ai-takes-center-stage-in-animations-future-top-tools-of-2024-revealed">https://opentools.ai/news/ai-takes-center-stage-in-animations-future-top-tools-of-2024-revealed</ext-link>.</mixed-citation></ref>
<ref id="ref-355"><label>[355]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Team SP</collab></person-group>. <article-title>SoA survey reveals a third of translators and quarter of illustrators losing work to AI. The Society of Authors</article-title>; <year>2024</year> <comment>[cited 2025 May 20]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://societyofauthors.org/2024/04/11/soa-survey-reveals-a-third-of-translators-and-quarter-of-illustrators-losing-work-to-ai/">https://societyofauthors.org/2024/04/11/soa-survey-reveals-a-third-of-translators-and-quarter-of-illustrators-losing-work-to-ai/</ext-link>.</mixed-citation></ref>
<ref id="ref-356"><label>[356]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hall</surname> <given-names>J</given-names></string-name>, <string-name><surname>Schofield</surname> <given-names>D</given-names></string-name></person-group>. <article-title>The value of creativity: human produced art vs. AI-generated art</article-title>. <source>Art Design Rev</source>. <year>2024</year>;<volume>13</volume>(<issue>1</issue>):<fpage>65</fpage>&#x2013;<lpage>88</lpage>. doi:<pub-id pub-id-type="doi">10.4236/adr.2025.131005</pub-id>.</mixed-citation></ref>
<ref id="ref-357"><label>[357]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Diffey</surname> <given-names>H</given-names></string-name></person-group>. <article-title>ScreenRant. 2025 [cited 2025 Jun 4]. Amid Toei&#x2019;s AI Controversy, the Anime Industry Is Pushing Back: &#x201C;Aren&#x2019;t We Shooting Ourselves In the Foot?&#x201D;</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://screenrant.com/anime-ai-japan-backlash-toei-animation-industry-future/">https://screenrant.com/anime-ai-japan-backlash-toei-animation-industry-future/</ext-link>.</mixed-citation></ref>
<ref id="ref-358"><label>[358]</label><mixed-citation publication-type="other"><article-title>Economic potential of generative AI | McKinsey</article-title>. <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier">https://www.mckinsey.com/capabilities/mckinsey-digital/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier</ext-link>.</mixed-citation></ref>
<ref id="ref-359"><label>[359]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>AI Development Cost: A Comprehensive Overview for 2025</collab></person-group>. <article-title>Prismetric</article-title>; <year>2025</year> <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.prismetric.com/ai-development-cost/">https://www.prismetric.com/ai-development-cost/</ext-link>.</mixed-citation></ref>
<ref id="ref-360"><label>[360]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Outshift by Cisco</collab></person-group>. <article-title>Outshift | AI infrastructure: Prepare your organization for transformation</article-title>. <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://outshift.com/blog/ai-infrastructure-how-to-prepare-your-organization-for-transformation">https://outshift.com/blog/ai-infrastructure-how-to-prepare-your-organization-for-transformation</ext-link>.</mixed-citation></ref>
<ref id="ref-361"><label>[361]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Animation World Network</collab></person-group>. <article-title>The future unscripted: the impact of generative artificial intelligence on entertainment industry jobs</article-title>. <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.awn.com/tag/future-unscripted-impact-generative-artificial-intelligence-entertainment-industry-jobs">https://www.awn.com/tag/future-unscripted-impact-generative-artificial-intelligence-entertainment-industry-jobs</ext-link>.</mixed-citation></ref>
<ref id="ref-362"><label>[362]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wolff</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Educating artists about AI in animation. Animation Magazine</article-title>. <year>2025</year>. <comment>[cited 2025 May 21]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.animationmagazine.net/2025/03/educating-artists-about-ai-in-animation/">https://www.animationmagazine.net/2025/03/educating-artists-about-ai-in-animation/</ext-link>.</mixed-citation></ref>
<ref id="ref-363"><label>[363]</label><mixed-citation publication-type="other"><article-title>Fine-Arts-2022.pdf</article-title>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://coloradostatefair.com/wp-content/uploads/2022/05/Fine-Arts-2022.pdf">https://coloradostatefair.com/wp-content/uploads/2022/05/Fine-Arts-2022.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-364"><label>[364]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Christies</collab></person-group>. <article-title>Christies.com. JESSE WOOLSTON</article-title>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://onlineonly.christies.com/s/augmented-intelligence/jesse-woolston-22/250105">https://onlineonly.christies.com/s/augmented-intelligence/jesse-woolston-22/250105</ext-link>.</mixed-citation></ref>
<ref id="ref-365"><label>[365]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Christies</collab></person-group>. <article-title>Christies.com. HOLLY HERNDON (B. 1980) AND MAT DRYHURST (B. 1984)</article-title>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://onlineonly.christies.com/s/augmented-intelligence/holly-herndon-b-1980-mat-dryhurst-b-1984-3/249745">https://onlineonly.christies.com/s/augmented-intelligence/holly-herndon-b-1980-mat-dryhurst-b-1984-3/249745</ext-link>.</mixed-citation></ref>
<ref id="ref-366"><label>[366]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>pixiv</collab></person-group>. <article-title>Popular illustrations and manga tagged &#x201C;AI&#x201D;</article-title>. <comment>[cited 2025 Aug 17]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.pixiv.net/en/tags/AI">https://www.pixiv.net/en/tags/AI</ext-link>.</mixed-citation></ref>
<ref id="ref-367"><label>[367]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>ArtStation</collab></person-group>. <article-title>ArtStation-Explore</article-title>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://artstation.com/">https://artstation.com/</ext-link>.</mixed-citation></ref>
<ref id="ref-368"><label>[368]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Anime News Network. 2025</collab></person-group>. <article-title>Novelist otsuichi co-directs generaidoscope, omnibus film produced entirely with generative AI</article-title>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.animenewsnetwork.com/news/2024-07-13/novelist-otsuichi-co-directs-generaidoscope-omnibus-film-produced-entirely-with-generative-ai/.213069">https://www.animenewsnetwork.com/news/2024-07-13/novelist-otsuichi-co-directs-generaidoscope-omnibus-film-produced-entirely-with-generative-ai/.213069</ext-link>.</mixed-citation></ref>
<ref id="ref-369"><label>[369]</label><mixed-citation publication-type="other"><article-title><italic>Twins Hinahima</italic>. In: Wikipedia</article-title>; <year>2025</year> <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://en.wikipedia.org/w/index.php?title=Twins_Hinahima&#x0026;oldid=1284878942">https://en.wikipedia.org/w/index.php?title=Twins_Hinahima&#x0026;oldid=1284878942</ext-link>.</mixed-citation></ref>
<ref id="ref-370"><label>[370]</label><mixed-citation publication-type="other"><article-title>Deep Learning Super Sampling. In: Wikipedia</article-title>. <year>2025</year> <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://en.wikipedia.org/w/index.php?title=Deep_Learning_Super_Sampling&#x0026;oldid=1291331644">https://en.wikipedia.org/w/index.php?title=Deep_Learning_Super_Sampling&#x0026;oldid=1291331644</ext-link>.</mixed-citation></ref>
<ref id="ref-371"><label>[371]</label><mixed-citation publication-type="other"><article-title><italic>Fossilized Wonders</italic>. In: Wikipedia</article-title>. <year>2025</year>. <comment>[cited 2025 Jun 5]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://en.wikipedia.org/w/index.php?title=Fossilized_Wonders&#x0026;oldid=1291840646">https://en.wikipedia.org/w/index.php?title=Fossilized_Wonders&#x0026;oldid=1291840646</ext-link>.</mixed-citation></ref>
</ref-list>
</back></article>



























