<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">74388</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2025.074388</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Inverse Design of Composite Materials Based on Latent Space and Bayesian Optimization</article-title>
<alt-title alt-title-type="left-running-head">Inverse Design of Composite Materials Based on Latent Space and Bayesian Optimization</alt-title>
<alt-title alt-title-type="right-running-head">Inverse Design of Composite Materials Based on Latent Space and Bayesian Optimization</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Lyu</surname><given-names>Xianrui</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ren</surname><given-names>Xiaodan</given-names></name><email>rxdtj@tongji.edu.cn</email></contrib>
<aff id="aff-1"><institution>Department of Structural Engineering, College of Civil Engineering, Tongji University</institution>, <addr-line>Shanghai, 200092</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Xiaodan Ren. Email: <email>rxdtj@tongji.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>29</day><month>1</month><year>2026</year>
</pub-date>
<volume>146</volume>
<issue>1</issue>
<elocation-id>1</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>10</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>01</day>
<month>12</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_74388.pdf"></self-uri>
<abstract>
<p>Inverse design of advanced materials represents a pivotal challenge in materials science. Leveraging the latent space of Variational Autoencoders (VAEs) for material optimization has emerged as a significant advancement in the field of material inverse design. However, VAEs are inherently prone to generating blurred images, posing challenges for precise inverse design and microstructure manufacturing. While increasing the dimensionality of the VAE latent space can mitigate reconstruction blurriness to some extent, it simultaneously imposes a substantial burden on target optimization due to an excessively high search space. To address these limitations, this study adopts a Variational Autoencoder guided Conditional Diffusion Generative Model (VAE-CDGM) framework integrated with Bayesian optimization to achieve the inverse design of composite materials with targeted mechanical properties. The VAE-CDGM model synergizes the strengths of VAEs and Denoising Diffusion Probabilistic Models (DDPM), enabling the generation of high-quality, sharp images while preserving a manipulable latent space. To accommodate varying dimensional requirements of the latent space, two optimization strategies are proposed. When the latent space dimensionality is excessively high, SHapley Additive exPlanations (SHAP) sensitivity analysis is employed to identify critical latent features for optimization within a reduced subspace. Conversely, direct optimization is performed in the low-dimensional latent space of VAE-CDGM when dimensionality is modest. The results demonstrate that both strategies accurately achieve the targeted design of composite materials while circumventing the blurred reconstruction flaws of VAEs, which offers a novel pathway for the precise design of advanced materials.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Variational autoencoder</kwd>
<kwd>denoising diffusion generation model</kwd>
<kwd>composite materials</kwd>
<kwd>Bayesian optimization</kwd>
<kwd>SHapley Additive exPlanations</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>52478198</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Shanghai Municipal Commission of Science and Technology</funding-source>
<award-id>24510711500</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Composite materials, renowned for their high strength, toughness, and multifunctionality, have found extensive applications across aerospace, new energy vehicles, and high-end manufacturing sectors [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-4">4</xref>]. As quintessential multiphase materials, their macroscopic properties, such as mechanical strength, thermal conductivity, and fatigue resistance, are governed not only by the intrinsic attributes of the matrix and reinforcement but also by microstructural features, including reinforcement distribution density, size distribution, and spatial orientation [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. For instance, a regular arrangement of reinforcement within the matrix can significantly enhance the strain hardening rate and flow stress, while a gradient arrangement, particularly when aligned parallel to the applied load, effectively mitigates damage propagation. Conversely, clustered microstructures may induce stress concentrations and precipitate premature failure [<xref ref-type="bibr" rid="ref-7">7</xref>]. Consequently, the precise design of microstructures to achieve targeted performance constitutes a pivotal objective in advanced materials engineering [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>Inverse design, a critical pathway bridging performance targets and microstructure, aims to infer optimal microstructural configurations from predefined macroscopic properties, with its efficiency and accuracy directly influencing the development cycle and engineering value of novel materials. Traditional inverse design methodologies rely on microstructure characterization techniques [<xref ref-type="bibr" rid="ref-9">9</xref>], employing physical descriptors [<xref ref-type="bibr" rid="ref-10">10</xref>], such as fiber volume fraction, average diameter, and spacing distribution [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>], to reduce the high-dimensional microstructural space to a lower-dimensional representation for optimization. However, these descriptors are often constrained by empirically defined rules, exhibiting limitations such as restricted spatial dimensionality and incomplete information representation. For example, volume fraction alone fails to capture localized clustering features, resulting in the inability to perform precise microstructure inverse design. As material systems continue to increase in complexity, conventional inverse design methodologies face growing challenges in simultaneously capturing both global microstructural patterns and localized features, emerging as a significant bottleneck in enhancing inverse design precision.</p>
<p>In recent years, deep generative models have ushered in a transformative paradigm for the precise characterization and inverse design of microstructures [<xref ref-type="bibr" rid="ref-9">9</xref>]. The pixel-based representation of microstructures fundamentally constitutes a high-dimensional random field. Meanwhile, the core objective of generative models which is high-dimensional probability density estimation [<xref ref-type="bibr" rid="ref-15">15</xref>], aligns intrinsically with the representational needs of microstructures, enabling the capture of higher-order statistical properties thereby providing a more comprehensive depiction of their complex information. Among these models, the Variational Autoencoder (VAE) emerges as a classical approach [<xref ref-type="bibr" rid="ref-16">16</xref>], employing an encoder to map high-dimensional microstructures into a low-dimensional latent space and a decoder to reconstruct the structures, demonstrating robust dimensionality reduction and generative capabilities. VAE has been extensively applied in the microstructure reconstruction of composite materials, metallic alloys, sandstones, and other material systems [<xref ref-type="bibr" rid="ref-17">17</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>], offering efficient tools for material characterization and design.</p>
<p>Leveraging the latent space of VAE for optimization design has emerged as a frontier strategy for inverse materials design. Recent studies have demonstrated its applicability across diverse material systems, aiming to generate materials with desired properties through latent space exploration and optimization. For instance, Wang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] demonstrated that VAE latent spaces provide a meaningful distance metric for assessing shape similarity, enabling vector operations to adjust mechanical properties and manipulate complex microstructures. Lew and Buehler [<xref ref-type="bibr" rid="ref-21">21</xref>] focused on compliance optimization of cantilever designs, encoding cantilever structures into a 2D latent space using a VAE and employing a Long Short-Term Memory (LSTM) network to learn optimization trajectories within that space. Kusampudi and Diehl [<xref ref-type="bibr" rid="ref-22">22</xref>] trained a VAE to identify descriptors from synthetic dual-phase steel microstructures, utilizing Bayesian optimization to determine optimal descriptor combinations that yield microstructures with specific yield strengths and reduced damage initiation sensitivity. Xue et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] employed the VAE to compress the low-resolution (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>28</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>28</mml:mn></mml:math></inline-formula>) microstructures of composite materials into a 10-dimensional latent space. As this low-dimensional representation was sufficient to capture the essential features at this low resolution, it allowed for efficient integration with Bayesian optimization, thereby facilitating the inverse design of the macroscopic elastic modulus. Extending this approach to composite laminates, Sun et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] developed a generative VAE framework that efficiently tackles the one-to-many mapping challenge in layup design. By structuring the latent space to separate stiffness-related features from sequence-style characteristics, their model enables rapid generation of numerous non-conventional laminates meeting target mechanical and manufacturing constraints. Park et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] employed various optimization algorithms, such as the single-code modification algorithm, genetic algorithm, and stochastic algorithm, within the VAE latent space to enhance the local energy stability of spin configurations and optimize global physical quantities, including topological indices, magnetization strength, energy, and directional correlations.</p>
<p>Nevertheless, VAEs exhibit inherent limitations in microstructure generation. Their objective function primarily optimizes for latent distribution regularization and pixel-wise reconstruction error, which often leads to decoded microstructures with lost details, such as blurred edges [<xref ref-type="bibr" rid="ref-26">26</xref>]. Although increasing the dimensionality of the latent space can partially alleviate reconstruction inaccuracies, it does not fully eliminate the blurring effect, while substantially increasing computational cost and complicating downstream inverse design tasks. As noted by Peng et al. [<xref ref-type="bibr" rid="ref-27">27</xref>], selecting an appropriate latent dimension requires balancing reconstruction quality against the curse of dimensionality. An undersized latent space lacks sufficient representational capacity, resulting in high reconstruction loss, whereas an oversized space introduces diminishing returns in accuracy while exponentially expanding the search space. This trade-off is particularly consequential when VAEs are coupled with Bayesian optimization (BO). Since BO is typically applicable to dimensions below 20, the microstructures generated by VAEs within this range are often blurry and inadequate for precise design and manufacturing objectives. Therefore, maintaining a low latent space dimension without compromising microstructural fidelity is essential to fully leverage Bayesian optimization in VAE-based inverse design frameworks.</p>
<p>To overcome this bottleneck, our research group previously introduced a Variational Autoencoder guided Conditional Diffusion Generative Model (VAE-CDGM) tailored for material design [<xref ref-type="bibr" rid="ref-28">28</xref>], integrating the low-dimensional mapping capabilities of VAEs with the high-fidelity generative attributes of Denoising Diffusion Probabilistic Models (DDPM). This model uses the blurred image generated by the VAE as a conditional input to restructure the forward diffusion process of DDPM. It then employs multi-step denoising iterations to progressively refine the microstructural details. This approach thus achieves high-fidelity microstructure reconstruction while maintaining a continuous and controllable latent space.</p>
<p>For the inverse design of composite materials, this study builds upon the VAE-CDGM framework, balancing latent space manipulability with high-quality generation, and proposes a dual-strategy inverse design approach integrated with Bayesian optimization. When higher latent space dimensionality is required, SHapley Additive exPlanations (SHAP) sensitivity analysis is employed to identify the most impactful dimensions for performance, enabling Bayesian optimization within a reduced subspace. Conversely, for lower-dimensional spaces, direct Bayesian optimization is performed in VAE-CDGM&#x2019;s latent space. Both strategies utilize VAE-CDGM to map optimized latent vectors back to the pixel space, generating high-quality microstructure images providing guidance for subsequent manufacturing processes. The material design framework proposed in the study is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The subsequent sections are organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> elaborates the construction principles of the VAE-CDGM model; <xref ref-type="sec" rid="s3">Section 3</xref> validates the latent space properties and reconstruction quality of VAE-CDGM; <xref ref-type="sec" rid="s4">Section 4</xref> experimentally confirms the model&#x2019;s generative performance and the efficacy of the inverse design strategy; <xref ref-type="sec" rid="s5">Section 5</xref> summarizes the findings, discusses the method&#x2019;s advantages and future improvement directions.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Material design framework based on VAE-CDGM and Bayesian optimization</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>VAE-CDGM Introduction</title>
<sec id="s2_1">
<label>2.1</label>
<title>Introduction to VAE</title>
<p>The Variational Autoencoder, a generative model integrating probabilistic modeling with deep learning, was pioneered by Kingma and Welling [<xref ref-type="bibr" rid="ref-16">16</xref>]. Its core innovation lies in employing variational inference to establish a probabilistic mapping between high-dimensional data and a low-dimensional latent space, facilitating the representation, generation, and reconstruction of microstructure images. Distinct from traditional autoencoders [<xref ref-type="bibr" rid="ref-29">29</xref>], VAE introduces regularization in the latent space, endowing it with continuity and interpretability, which lays a theoretical foundation for interpolation-based generation and optimization-driven design in materials engineering.</p>
<p>The generative process assumes that observed data <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:math></inline-formula> is produced via <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x|z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where the latent variable <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:math></inline-formula> follows a prior distribution <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, typically a standard normal distribution <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The objective is to learn the joint distribution <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x,&#x00A0;z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x|z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to approximate the true data distribution. However, directly maximizing the log-likelihood <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x,&#x00A0;z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:math></inline-formula> in high-dimensional spaces is intractable due to the complex integral over the latent variables.</p>
<p>To address this, VAE employs variational inference and Jensen&#x2019;s inequality to avoid direct integration. By introducing an approximate posterior distribution, the marginal log-likelihood can be expressed as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mi>d</mml:mi><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>here, since <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is a strictly concave function, according to Jensen&#x2019;s inequality, the above expression can be transformed into,
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The expectation term on the right side of the above equation is the core of VAE, the Evidence Lower Bound (ELBO),
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mtext>ELBO</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>By applying Jensen&#x2019;s inequality, the intractable log-likelihood <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is transformed into a computationally tractable lower bound, the ELBO. Since <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mtext>ELBO</mml:mtext></mml:math></inline-formula>, maximizing the ELBO indirectly approximates the maximization of <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Splitting the logarithmic terms in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> and substituting them into the expectation-based definition of the ELBO yields:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>ELBO</mml:mtext></mml:mrow></mml:mtd><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd /><mml:mtd><mml:mi></mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Since machine learning typically minimizes loss functions via gradient descent, while the objective of VAE is to maximize the ELBO, this is equivalent to minimizing the negative ELBO (&#x2212;ELBO). The VAE optimization objective can ultimately be written in the following form,
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>V</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>ELBO</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>log</mml:mtext></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x|z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the conditional likelihood modeled by the decoder, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z|x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the approximate posterior modeled by the encoder, and <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>&#x03D5;</mml:mi></mml:math></inline-formula> denote the parameters of the decoder and encoder, respectively. The first term represents the reconstruction error, quantifying the similarity between generated and original data, while the second term serves as a regularization constraint, ensuring the approximate posterior aligns with the prior distribution. The schematic diagram of the VAE framework is shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Schematic diagram of VAE architecture</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-2.tif"/>
</fig>
<p>In materials science, VAE&#x2019;s robust non-linear dimensionality reduction and generative capabilities have been widely leveraged for microstructure characterization and reconstruction across diverse material systems. Its probabilistic framework enables interpolation and optimization operations within the continuous latent space, providing a foundation for target-oriented material design. Despite these advantages, VAE-generated outputs often suffer from blurriness due to the loss&#x2019;s emphasis on mean-field approximations. Moreover, the choice of latent space dimensionality critically influences both generation quality and computational complexity. Low-dimensional spaces may result in the loss of essential microstructural features, whereas high-dimensional spaces, while potentially reducing blurriness, significantly increase the computational burden and optimization challenges in subsequent inverse design tasks.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Introduction to Denoising Diffusion Probabilistic Model (DDPM)</title>
<p>The Denoising Diffusion Probabilistic Model (DDPM), a deep learning framework rooted in probabilistic generative processes, was initially proposed by Sohl-Dickstein et al. in 2015 [<xref ref-type="bibr" rid="ref-30">30</xref>] and subsequently developed by Ho et al. in 2020 [<xref ref-type="bibr" rid="ref-31">31</xref>]. DDPM achieves high-precision modeling of complex high-dimensional data distributions by simulating a stepwise noising and denoising process. Compared to VAEs, DDPM excels in generating high-quality samples, demonstrating significant advantages in image generation and microstructure reconstruction [<xref ref-type="bibr" rid="ref-32">32</xref>&#x2013;<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>The core concept of DDPM revolves around a forward diffusion process and a reverse denoising process to model the data probability distribution, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The forward process incrementally adds Gaussian noise to the original data <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, transitioning it over T steps toward an isotropic Gaussian distribution. This process is defined as a Markov chain:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> represents the noise scheduling parameter, controlling the intensity of noise addition. According to the properties of Gaussian distribution, the forward process can be directly expressed as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:msqrt><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>:=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>:=</mml:mo><mml:munderover><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>t</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The forward and reverse processes of diffusion model</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-3.tif"/>
</fig>
<p>The reverse process aims to recover the original data from noise,
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the mean and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> is a predefined or learned variance. According to Bayes&#x2019; theorem, the mean <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> can be derived as:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:msub><mml:mi mathvariant="bold-italic">&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The objective of DDPM is to minimize the KL divergence between the true posterior distribution <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and the model-predicted distribution <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>|</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Moreover, since,
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mo mathvariant="bold" stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mi mathvariant="bold-italic">&#x03F5;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi mathvariant="bold-italic">&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:msub><mml:mi mathvariant="bold-italic">&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Substituting the above equation into the optimization objective, the objective becomes:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mfrac><mml:msubsup><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mn>2</mml:mn><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Removing the coefficient yields the final loss function of DDPM,
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>simple</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> is the noise added during the forward process, and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the noise predicted by the neural network.</p>
<p>Compared to the blurred images generated by VAEs, DDPM&#x2019;s multi-step denoising process markedly enhances detail fidelity, providing reliable samples for material performance prediction and reverse design. However, the diffusion model does not undergo dimensionality reduction during implementation, resulting in its latent space and pixel space being the same dimensions, making it impossible to use it as a search space for target optimization.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Introduction to VAE-CDGM</title>
<p>The preceding analyses highlight that while VAEs offer a low-dimensional, continuous latent space conducive to optimization, they are prone to producing blurred reconstructions. Conversely, DDPM excels in high-quality image reconstruction but lacks a manipulable latent space for material optimization. To address these complementary limitations, the VAE-CDGM integrates the strengths of VAEs and DDPMs, ensuring the presence of a continuous latent space while enhancing microstructure reconstruction quality [<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<p>Unlike standard DDPM, VAE-CDGM incorporates a conditional diffusion process during the forward process:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:math></inline-formula> is the latent vector from the VAE, and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the blurred image obtained by mapping <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:math></inline-formula> through the VAE decoder. In contrast to <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>, the VAE-CDGM diffusion process does not fully transition the clean image to pure Gaussian noise but applies a mean correction, shifting the zero-mean noise toward the blurred image&#x2019;s mean, as illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Mean correction process of VAE-CDGM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-4.tif"/>
</fig>
<p>Subsequently, based on Bayesian inference, the reverse process can be expressed as:
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math></disp-formula>with the posterior mean and variance derived from the forward process as:
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msqrt><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:msqrt></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B2;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. The loss function for VAE-CDGM is thus formulated as:
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mover><mml:mi>&#x03BC;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo>|</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msqrt><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:msqrt><mml:mi>&#x03F5;</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mo>|</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>Traditional conditional diffusion models typically employ a standard diffusion process, incrementally adding Gaussian noise to images until they degenerate into pure noise, with conditional information (e.g., low-quality images) appended as input to the noise prediction network. In contrast, VAE-CDGM explicitly models the image degradation process, progressively degrading high-quality images <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:math></inline-formula> into low-quality images <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:math></inline-formula> with superimposed noise. The reverse process of traditional models initiates from pure noise, relying on conditional information to incrementally reconstruct content. However, the high initial randomness often leads to deviations from the true content. VAE-CDGM&#x2019;s reverse process starts from the noisy low-quality image, directly optimizing the recovery path. This approach avoids expending model capacity to regenerate existing features from pure noise, initializing the reverse process closer to the target distribution and reducing generative bias.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Model Architecture Design</title>
<p>The VAE-CDGM model adopts a two-phase training strategy [<xref ref-type="bibr" rid="ref-35">35</xref>], fully decoupling the training parameters of its constituent components. In the first phase, the model focuses on training the VAE component, where the encoder network maps input microstructure images <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:math></inline-formula> into low-dimensional latent variables <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow></mml:math></inline-formula>, and the decoder network reconstructs these latent variables into blurred microstructure images <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:math></inline-formula>. The primary optimization objective during this phase is to minimize the loss function <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Upon completion of the VAE training, its parameters are fully frozen, and the model transitions to the second phase, concentrating on training the conditional diffusion model. This phase utilizes the blurred reconstructions from the first phase as conditional inputs, employing a meticulously designed Markov chain process. In the forward process, Gaussian noise is incrementally added to the original data, while the reverse process learns a complex denoising process to generate high-fidelity microstructure details. The primary optimization objective during this phase is to minimize the loss function <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>D</mml:mi><mml:mi>P</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The overall workflow is depicted in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Schematic diagram of VAE-CDGM framework</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-5.tif"/>
</fig>
<p>The model&#x2019;s first branch comprises a convolutional VAE network. The encoder employs a five-layer downsampling module, each consisting of a Conv2d, BatchNorm2d, and ReLU combination, compressing a <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula> input image into a <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mn>4</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula> feature map, which is subsequently mapped to a low-dimensional latent space. The corresponding decoder utilizes a symmetric five-layer transposed convolutional structure (ConvTranspose2d &#x002B; BatchNorm2d &#x002B; ReLU) to reconstruct the latent variables into blurred microstructure images. The second branch features a U-Net-based conditional diffusion model, incorporating 16 residual blocks and multiple attention layers. This branch takes the VAE-output blurred reconstructions as conditional inputs, employing downsampling and upsampling paths for feature extraction and fusion. The specific architectural configuration of the U-Net is detailed in <xref ref-type="table" rid="table-1">Table 1</xref>. The model comprises approximately 20.1 million trainable parameters. Training was conducted for 100 epochs on an NVIDIA GeForce320 RTX 4070 Ti GPU, with a total training time of approximately 105,767 s.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>U-Net network architecture configuration</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Stage</th>
<th>Resolution</th>
<th>Channels</th>
<th>Key operations</th>
<th>Down/Up-sampling method</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input</td>
<td><inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mn>1</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>64</mml:mn></mml:math></inline-formula></td>
<td>Conditional concatenation &#x002B; initial convolution</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Encoder</td>
<td><inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>8</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mn>64</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1024</mml:mn></mml:math></inline-formula></td>
<td>Downsampling &#x002B; residual blocks</td>
<td>Average pooling</td>
</tr>
<tr>
<td>Bottleneck</td>
<td><inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mn>8</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>8</mml:mn></mml:math></inline-formula></td>
<td>1024</td>
<td>Feature extraction</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Decoder</td>
<td><inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mn>8</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>8</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mn>1024</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>64</mml:mn></mml:math></inline-formula></td>
<td>Upsampling &#x002B; skip connections</td>
<td>Nearest neighbor interpolation</td>
</tr>
<tr>
<td>Output</td>
<td><inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mn>64</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula></td>
<td>Final convolution</td>
<td>&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Characterization and Reconstruction of Composites Based on VAE-CDGM</title>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset Generation</title>
<p>This study generates two-dimensional microstructure images of composite materials with randomly distributed circular inclusions to simulate diverse inclusion patterns. The images are synthesized at a resolution of <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula> pixels, with inclusion radii varying randomly between 10 and 15 pixels, and 12 to 18 circular inclusions randomly placed per image. The center coordinates <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are randomly sampled within the image bounds <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>x</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>r</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>r</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> to prevent edge truncation, where <italic>r</italic> is the radius, and <italic>W</italic> and <italic>H</italic> are the image width and height, respectively. A total of 11,893 composite microstructure images were generated for training the VAE-CDGM model. Both the VAE and DDPM components were configured with a batch size of 64, optimized using the Adam optimizer with a learning rate of 0.001.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Generation Quality Validation</title>
<p>In this study, we conducted a comparative analysis of the microstructure generation quality for composite materials with circular inclusions, using VAE and the proposed VAE-CDGM model at latent space dimensions of 256 and 16, as sown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. For the 256-dimensional latent space, the VAE generated microstructures with improved detail retention compared to its 16-dimensional counterpart, capturing finer geometric features such as inclusion boundaries and spatial distributions. However, the VAE outputs still exhibited noticeable blurriness, particularly in boundary regularity and inclusion size accuracy. In contrast, the VAE-CDGM model at 256 dimensions significantly enhanced generation quality, producing microstructures with sharper boundaries and higher fidelity. When reducing the latent space to 16 dimensions, the VAE suffered from substantial information loss, resulting in overly blurry microstructures with diminished topological accuracy. Conversely, the VAE-CDGM model maintained superior generation quality even at 16 dimensions, leveraging the iterative denoising process of DDPM to reconstruct detailed microstructures while preserving the low-dimensional latent representation of VAE. These results underscore the VAE-CDGM model&#x2019;s ability to balance latent space dimensionality and high-fidelity microstructure generation, making it a robust tool for inverse design applications in composite materials.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparison of VAE and VAE-CDGM generation quality under different latent spaces (<bold>a</bold>) VAE-256; (<bold>b</bold>) VAE-CDGM-256; (<bold>c</bold>) VAE-16; (<bold>d</bold>) VAE-CDGM-16</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-6a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-6b.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Latent Space Analysis</title>
<p>To gain deeper insights into the latent space properties of the VAE-CDGM model, we conducted a detailed analysis of the 16-dimensional latent space. To facilitate visualization, Principal Component Analysis (PCA) was employed to reduce the dimensionality to three principal components, enabling the projection of the latent space onto the first two principal components (PC1 and PC2). The label assigned to each data point represents the homogenized stress of the composite material under uniaxial compression. The finite element model configuration, including mesh discretization and boundary conditions, is depicted in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. Each microstructure image was directly converted into a corresponding finite element model using the python script that maps pixel intensities to material properties, with each pixel corresponding to one CPS4R finite element. The model simulates the mechanical behavior of the composite using two linearly elastic materials: Material- matrix with a Young&#x2019;s modulus <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>73</mml:mn><mml:mspace width="thinmathspace" /><mml:mtext>GPa</mml:mtext></mml:math></inline-formula> and Poisson&#x2019;s ratio <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>&#x03BD;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.33</mml:mn></mml:math></inline-formula>, and Material- reinforcement with <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>420</mml:mn><mml:mspace width="thinmathspace" /><mml:mtext>GPa</mml:mtext></mml:math></inline-formula> and <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03BD;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.13</mml:mn></mml:math></inline-formula>. A vertical displacement load of 5 mm was applied, and the quasi-static response was solved using a static analysis step.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Microstructure image and finite element setup of composite material (<bold>a</bold>) Microstructure image of composite material; (<bold>b</bold>) Mesh discretization and boundary setting</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-7.tif"/>
</fig>
<p>The PCA visualization results, presented in <xref ref-type="fig" rid="fig-8">Fig. 8a</xref>, reveal that the 16-dimensional latent space retains significant structural characteristics post-reduction. In stark contrast to the diffusion model&#x2019;s latent space, which exhibits a disordered and unstructured distribution under identical PCA conditions (<xref ref-type="fig" rid="fig-8">Fig. 8b</xref>), the homogenized stress exhibits a pronounced clustered distribution within the PC1-PC2 plane. Low-stress microstructures concentrated in the lower-left region and high-stress microstructures in the upper-right region, accompanied by a smooth gradient along the diagonal. This clustering in the PCA space indicates that the VAE-CDGM model successfully compresses high-dimensional microstructural information into a low-dimensional representation while preserving critical features correlated with mechanical performance. The smooth gradient further reflects the continuity of the latent space, enabling interpolation or optimization to explore performance variations, which is conducive to efficient inverse design. This structured latent space not only validates the superior representational capability of the VAE-CDGM model but also provides a reliable design space for target-oriented material optimization.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Distribution characteristics of composite materials in (<bold>a</bold>) VAE-CDGM latent space and (<bold>b</bold>) DDPM latent space</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Inverse Design of Composites Based on Bayesian Optimization</title>
<sec id="s4_1">
<label>4.1</label>
<title>Introduction to Bayesian Optimization</title>
<p>Bayesian Optimization (BO) represents an efficient global optimization methodology that constructs a probabilistic surrogate model of the target function and leverages an acquisition function to guide the search process, thereby significantly reducing the number of evaluations required for parameter optimization. In the realm of materials science, BO has been extensively adopted for inverse design applications [<xref ref-type="bibr" rid="ref-36">36</xref>&#x2013;<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
<p>The primary objective of Bayesian optimization is to identify the global optimum of a target function <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> within its defined domain <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> typically represents a computationally expensive black-box function lacking an analytical form or gradient information [<xref ref-type="bibr" rid="ref-41">41</xref>]. This method iteratively proceeds through two fundamental steps to achieve efficient optimization. First, a probabilistic surrogate model is constructed based on existing observational data <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, with the Gaussian Process (GP) being the most commonly employed choice due to its ability to quantify uncertainty [<xref ref-type="bibr" rid="ref-42">42</xref>]. The GP assumes that the target function follows a multivariate normal distribution, where the kernel function governs the smoothness and correlation structure of the function across the domain. Given the observational data, the GP yields a posterior distribution over <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, expressed as <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> are the predictive mean and variance, respectively, computed from the kernel function and data, providing a comprehensive characterization of the function&#x2019;s uncertainty and expected behavior. Second, an acquisition function <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is utilized to select the next evaluation point, strategically balancing exploration, sampling regions with high uncertainty, and exploitation, sampling regions with high predicted function values to accelerate convergence toward the optimum.</p>
<p>The Expected Improvement (EI) acquisition function stands as a widely adopted choice in Bayesian optimization, renowned for its effective balance between exploration and exploitation [<xref ref-type="bibr" rid="ref-43">43</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>]. EI quantifies the anticipated improvement in the objective function <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> at a given point <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow></mml:math></inline-formula>, defined as:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>EI</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> denotes the current best-known function value. Assuming a Gaussian Process (GP) posterior distribution <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, EI can be computed analytically as:
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>EI</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>z</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represent the cumulative distribution function and probability density function of the standard normal distribution, respectively. The first term promotes exploitation by favoring points with high predicted values, while the second term encourages exploration by prioritizing regions with significant uncertainty. Each iteration of the Bayesian optimization loop selects the next point by maximizing the acquisition function:
<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>next</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>arg</mml:mtext></mml:mrow><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mrow><mml:mtext>EI</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">x</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This process iterates iteratively until the predetermined number of evaluations or convergence conditions are reached. However, the computational complexity of GP regression, coupled with the sparsity of data in high-dimensional spaces, poses significant challenges, typically constraining its effectiveness to low-dimensional problems (<inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>d</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula>).</p>
<p>To address these limitations, the proposed VAE-CDGM model integrates the low-dimensional representation capabilities of VAE with the high-fidelity image generation prowess of DDPM, enabling the generation of detailed microstructures for composite materials while maintaining an efficient latent space. Theoretically, the model&#x2019;s capacity to produce high-quality microstructure images from low-dimensional latent spaces (&#x003C;20 dimensions) suggests that Bayesian optimization can be directly applied to such spaces to efficiently search for microstructures with target mechanical performance. Nevertheless, to reconcile the practical trade-offs among representational completeness, computational efficiency, and optimization robustness, we devised two tailored optimization strategies. For high-dimensional latent spaces (dimensionality &#x003E; 20), we conducted SHAP analysis using a fully connected neural network (FCN) to quantify feature importance, subsequently selecting the top 16 most influential dimensions according to their mean absolute SHAP values. This dimension reduction ensures that Bayesian optimization, guided by the EI acquisition function, operates within a compact yet informative subspace, enhancing convergence speed and robustness. Conversely, for low-dimensional latent spaces (&#x003C;20 dimensions), the VAE-CDGM model provides a concise representation, enabling direct Bayesian optimization without further reduction. By proposing this dual-strategy framework, we address the distinct computational and representational demands of high- and low-dimensional latent spaces, harnessing the VAE-CDGM to facilitate efficient and robust inverse design of composite microstructures.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Bayesian Optimization under 256 Dimensional Latent Space</title>
<p>The VAE-CDGM framework employs a decoupled training strategy for its VAE and DDPM components, which enhances compatibility with pre-trained VAE models across diverse tasks. This architecture offers significant flexibility by enabling the integration of existing VAEs without requiring computationally expensive retraining. However, when merging pre-trained VAEs for design tasks, especially those with high-dimensional latent spaces where retraining is prohibitively costly, a critical challenge is the efficient optimization of targets in such high-dimensional spaces. This challenge is further compounded by stringent requirements for reconstruction fidelity of VAE or the need to model intricate high-order statistical correlations. However, direct optimization in such high-dimensional spaces introduces substantial challenges. To mitigate these issues, we employed SHAP [<xref ref-type="bibr" rid="ref-45">45</xref>] sensitivity analysis to identify the 16 most influential dimensions based on mean absolute SHAP values, enabling Bayesian optimization to operate within a compact yet information-rich subspace that preserves critical features while significantly reducing computational overhead.</p>
<p>To perform sensitivity analysis and identify critical latent features governing mechanical properties, we first developed a surrogate mapping model between the latent space and material performance. This approach is fundamentally similar to CNN-based methods [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>], which also operate by extracting abstract low-dimensional features from high-dimensional images and then mapping them to physical properties via a fully connected neural network. In the present study, the VAE encoder replaces the CNN as the feature extractor. Specifically, we constructed a FCN to predict homogenized stress from 256-dimensional latent vectors. The network architecture comprised an input layer (256 dimensions), followed by hidden layers with dimensions of 512, 256, 128, and 64, each employing ReLU activation functions, and a final single-output linear layer. The model was trained on a dataset of 11,893 samples, randomly partitioned into an 80% training set (9514 samples) and a 20% validation set (2379 samples). Training was performed using the Adam optimizer with a learning rate of 0.001 and mean squared error (MSE) loss. The selection of the above hyperparameters was determined through ablation experiments and trial and error. Meanwhile, we added dropout layers (<italic>p</italic> &#x003D; 0.1) after each hidden layer. Early stopping was enforced with a patience of 10 epochs and a minimum validation improvement of <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.001</mml:mn></mml:math></inline-formula>. Under this protocol, the model achieves <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msup><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>0.97</mml:mn></mml:math></inline-formula> on the training set and <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msup><mml:mi>R</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:mn>0.69</mml:mn></mml:math></inline-formula> on the validation set after convergence, as shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. The significant discrepancy between the training and validation performance indicates a degree of overfitting. Nonetheless, the model is still able to identify correlations between latent features and mechanical properties, which serves as a reasonable reference for subspace dimensionality reduction [<xref ref-type="bibr" rid="ref-22">22</xref>]. Therefore, in this study, we directly utilized the 16 relatively important features identified through SHAP analysis, without further expanding the dataset for training.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Performance of the surrogate model on the training and testing sets</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-9.tif"/>
</fig>
<p>To identify the latent dimensions most critical for homogenized stress prediction, we conducted a comprehensive sensitivity analysis using SHAP implemented through the DeepExplainer. SHAP is a framework rooted in cooperative game theory, specifically the Shapley Value, which aims to fairly allocate the prediction output for a single instance to each of its features. The SHAP value <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> for a feature is calculated according to the following equation:
<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mo>&#x2286;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>i</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:mrow><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>!</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>!</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo>!</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>S</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>i</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p>here, <italic>N</italic> is the set of all features, <italic>S</italic> is a subset of features excluding <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>i</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>S</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the model&#x2019;s prediction using the subset of features <italic>S</italic>. This formula evaluates the marginal contribution of feature <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>i</mml:mi></mml:math></inline-formula> across all possible subsets of features, weighted by the number of ways each subset can be formed.</p>
<p>This approach quantitatively evaluated the contribution of each latent dimension to the predictions generated by our FCN by computing SHAP values. The mean absolute SHAP values were ranked, and the top 16 dimensions were selected, as they captured the majority of the predictive influence. To ensure the stability and reliability of the SHAP analysis, this study performed multiple random splits of the dataset (3 splits <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 2 repeats, resulting in 6 independent subdatasets) for comprehensive validation. The top 16 important features remained identical across all 6 independent analyses, as shown in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. Furthermore, all 16 features exhibit coefficient of variation (CV) values below 0.1, confirming that each feature&#x2019;s importance remains highly stable under data resampling. The most influential feature is Feature 34, as indicated by its highest mean SHAP value. Through SHAP dependency analysis, this study systematically investigated the influence mechanisms of different latent dimension features on the prediction of homogenized stress. Specifically, <xref ref-type="fig" rid="fig-11">Fig. 11</xref> presents dependency plots that reveal the relationships between the six most important latent variables and SHAP value. The results demonstrate that Feature 34 exhibits a significant negative correlation with homogenized stress. When this feature value is negative and decreases, its positive contribution to homogenized stress significantly enhances; whereas when the feature value is positive and increases, it demonstrates a substantial weakening effect on homogenized stress. Notably, Features 151, 7, 213, 47, and 38 display distinct U-shaped relationships with homogenized stress. When these feature values are below the critical value (zero), SHAP values decrease monotonically with increasing feature values, indicating a negative correlation with homogenized stress. When feature values exceed the critical threshold, SHAP values increase monotonically with feature values, transitioning to a positive correlation. Further analysis reveals that the overall influence of Features 34 and 151 is predominantly negative, while Features 213, 47, and 38 exhibit primarily positive characteristic responses. The corresponding microstructural variations resulting from systematic perturbations of these key latent dimensions are visualized in Appendix A (see <xref ref-type="fig" rid="fig-16">Fig. A1</xref>). This quantitative feature-response relationship provides important insights for latent space-based optimization design: through targeted adjustment of specific latent dimension values, precise control of homogenized stress can be achieved, offering theoretical guidance for optimizing material performance design.</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Top 16 feature importance ranking based on mean absolute SHAP values</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-10.tif"/>
</fig><fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>SHAP dependency plots for homogenized stress predictions (<bold>a</bold>) Feature 34; (<bold>b</bold>) Feature 151; (<bold>c</bold>) Feature 7; (<bold>d</bold>) Feature 213; (<bold>e</bold>) Feature 47; (<bold>f</bold>) Feature 38</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-11a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-11b.tif"/>
</fig>
<p>Subsequently, Bayesian optimization was performed within the 16-dimensional subspace to identify latent vectors corresponding to a target homogenized stress of <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mn>4.2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa under uniaxial compression. It is noteworthy that this target value was deliberately set beyond the maximum homogenized stress of approximately <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mn>3.9</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa observed in the dataset, explicitly testing the model&#x2019;s capability to extrapolate beyond the known design space. Employing a Gaussian Process (GP) as the surrogate model and the EI acquisition function, the optimization iteratively navigated the subspace to maximize the likelihood of achieving the target homogenized stress. The GP was initialized with 6 random samples of latent vectors, and the EI function balanced exploration of regions with high uncertainty against exploitation of areas with elevated predicted performance. The optimization procedure maintained the remaining 240 latent dimensions as fixed parameters while actively exploring combinations within the critical 16-dimensional subspace. At each iteration, five candidate 16-dimensional vectors were generated and assembled with a fixed 240-dimensional vector to form complete latent representations, which were subsequently decoded into microstructure images by the VAE decoder. These microstructures were evaluated via finite element analysis (FEA) to determine their mechanical performance, with the results incorporated into the GP model to refine subsequent predictions. To validate the stability and reliability of the optimization process, three independent optimization trials were conducted for the same target. As depicted in <xref ref-type="fig" rid="fig-12">Fig. 12a</xref>, the target function value, plotted as the mean homogenized stress per iteration, exhibited rapid convergence after only 4&#x2013;6 iterations. During the final iteration and evaluation, the microstructure image that was blurry yet closest to the target was identified. It was then refined into a high-quality microstructure image using the VAE-CDGM model. Notably, to significantly reduce the computational cost associated with the diffusion sampling process, we employed the Denoising Diffusion Implicit Model (DDIM) during the denoising stage [<xref ref-type="bibr" rid="ref-48">48</xref>]. The microstructural changes during the iteration process are shown in <xref ref-type="fig" rid="fig-12">Fig. 12b</xref>. Comparative analysis of the decoded microstructures presented in <xref ref-type="fig" rid="fig-13">Fig. 13</xref> revealed significant differences in generation quality. While VAE-produced images exhibited blurred inclusion boundaries and structural artifacts, VAE-CDGM generated microstructures with sharply defined, geometrically regular inclusions. FEA quantification of mechanical performance yielded homogenized stress of <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mn>4.19</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa for VAE-generated structures and <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mn>4.10</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa for VAE-CDGM-generated structures on average over three independent optimization processes, both approximating the target of <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mn>4.2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa. The slight discrepancy arises from the fundamental difference in structural fidelity. The marginally higher stress value in the VAE-generated structure is attributed to the numerous irregular, fine-scale inclusions inherent in its blurred output, which create localized stress concentrations and thereby inflate the homogenized stress. In contrast, the refinement process of VAE-CDGM eliminates these geometrical irregularities, resulting in a more uniform stress distribution and a slightly lower, yet potentially more physically realistic, homogenized stress value. Although the VAE-CDGM result showed a minor deviation, this was considered operationally acceptable given its superior microstructural fidelity, a critical factor for manufacturability. The demonstrated capability to maintain mechanical performance while significantly improving microstructure image quality highlights the practical advantage of the VAE-CDGM framework for materials design applications.</p>
<fig id="fig-12">
<label>Figure 12</label>
<caption>
<title>Bayesian optimization iteration process (<bold>a</bold>) Homogenized stress changes during the iteration process; (<bold>b</bold>) Microstructure changes during the iterative process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-12.tif"/>
</fig><fig id="fig-13">
<label>Figure 13</label>
<caption>
<title>Bayesian optimization results: (<bold>a</bold>,<bold>c</bold>,<bold>e</bold>) VAE decoded microstructure; (<bold>b</bold>,<bold>d</bold>,<bold>f</bold>) FEM verification of VAE; (<bold>g</bold>,<bold>i</bold>,<bold>k</bold>) VAE-CDGM decoded microstructure; (<bold>h</bold>,<bold>j</bold>,<bold>l</bold>) FEM verification of VAE-CDGM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-13.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Bayesian Optimization under 16 Dimensional Latent Space</title>
<p>A 16-dimensional latent space is preferentially selected when computational efficiency and rapid optimization are prioritized, or when prior analyses demonstrate that a compact representation adequately captures the key microstructural features and statistical information. This latent space is derived from the VAE-CDGM model, where prior PCA visualization reveals a structured distribution characterized by distinct clusters of homogenized stress values along smooth gradients. This organized topology, coupled with the VAE-CDGM model&#x2019;s capacity to generate high-fidelity microstructure images supports direct optimization without necessitating additional dimensionality reduction. A Gaussian Process (GP) surrogate model was initialized using a random sample of latent vectors from the 16-dimensional space, with the EI acquisition function guiding the search process. The EI function strategically balances exploration of regions with high predictive uncertainty and exploitation of areas with superior predicted performance, iteratively optimizing the latent vectors to achieve the target homogenized stress of <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mn>2.5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa, which was deliberately chosen to be lower than the minimum value in the dataset.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-14">Fig. 14</xref>, Bayesian optimization achieved convergence to the target performance after merely 2 iterations, requiring only 10 finite element evaluations, demonstrating remarkable efficiency in the optimization process. The optimized 16-dimensional latent vector was subsequently decoded into realistic microstructure images using both the VAE and VAE-CDGM models, with the results presented in <xref ref-type="fig" rid="fig-15">Fig. 15</xref>. The VAE-decoded images, constrained by significant information loss inherent to the 16-dimensional representation, exhibited blurred features with indistinct circular inclusions, limiting their fidelity to the target microstructure. In contrast, the VAE-CDGM-decoded images displayed enhanced clarity and detailed characteristics, including well-defined circular inclusions, reflecting its superior generative capacity. FEA quantification of mechanical performance yielded homogenized stress of <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mn>2.52</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa for VAE-generated structures and <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mn>2.51</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa for VAE-CDGM-generated structures on average over three independent optimization processes, both approximating the target of <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mn>2.5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> MPa. These results underscore that, within a low-dimensional latent space (d &#x003C; 20), the VAE-CDGM model not only facilitates rapid and effective Bayesian optimization but also ensures high-quality decoded microstructure images. This dual capability holds significant implications for guiding the design and manufacturing of advanced composite materials.</p>
<fig id="fig-14">
<label>Figure 14</label>
<caption>
<title>Bayesian optimization iteration process (<bold>a</bold>) Homogenized stress changes during the iteration process; (<bold>b</bold>) Microstructure changes during the iterative process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-14.tif"/>
</fig><fig id="fig-15">
<label>Figure 15</label>
<caption>
<title>Bayesian optimization results: (<bold>a</bold>,<bold>c</bold>,<bold>e</bold>) VAE decoded microstructure; (<bold>b</bold>,<bold>d</bold>,<bold>f</bold>) FEM verification of VAE; (<bold>g</bold>,<bold>i</bold>,<bold>k</bold>) VAE-CDGM decoded microstructure; (<bold>h</bold>,<bold>j</bold>,<bold>l</bold>) FEM verification of VAE-CDGM</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-15.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This study proposes a inverse design framework for composite materials that integrates the VAE-CDGM with Bayesian optimization. The VAE-CDGM model addresses the blurriness of VAE reconstructions while providing a continuous and differentiable latent space. Bayesian optimization, leveraging a probabilistic surrogate model and acquisition function, achieves convergence to target performance with minimal finite element evaluations. To accommodate varying dimensional requirements of the latent space, two distinct optimization strategies are introduced. For high-dimensional latent spaces, a fully connected neural network was developed to model the relationship between input features and target properties. Subsequently, SHapley Additive exPlanations (SHAP) analysis was employed to quantify the contribution of individual input features to the model&#x2019;s predictions, identifying key factors influencing target performance, and Bayesian optimization was performed within the subspace defined by these critical features. For low-dimensional latent spaces, optimization was conducted directly within the design space provided by VAE-CDGM. In the case study targeting the mechanical property of composite materials, convergence to the target performance was achieved with only 2&#x2013;6 iterations, corresponding to 10&#x2013;30 finite element analyses, respectively. Notably, microstructures decoded by VAE-CDGM exhibited superior clarity and fidelity compared to those from VAE. Collectively, this inverse design framework achieves both high efficiency and precise microstructure reconstruction, offering an innovative and viable pathway for the inverse design of advanced materials.</p>
<p>It is noteworthy that in this study, while VAE-CDGM refines the blurred images generated by the VAE, it also tends to mitigate stress concentrations induced by fine inclusions present in the VAE-generated blurry images. As a result, the homogenized stress of the microstructures refined by VAE-CDGM may be slightly lower than the target stress. Although this outcome is both predictable and acceptable, future work could explore performing VAE-CDGM refinement on the decoded VAE images at each iteration to address this discrepancy. Meanwhile, while the selection of key latent dimensions through SHAP analysis enables effective integration with Bayesian optimization, it does entail a partial loss of the full latent space&#x2019;s representational capacity for microstructures. Future work will explore hierarchical optimization strategies, commencing with preliminary exploration in low-dimensional spaces before progressively expanding to refined searches in higher dimensions. Additionally, the integration of Bayesian optimization with diffusion models within the latent space represents a promising direction for future research.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The financial support provided by National Natural Science Foundation of China (Grant No. 52478198) and Shanghai Municipal Commission of Science and Technology (Grant No. 24510711500) is greatly appreciated.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, methodology and writing&#x2014;original draft preparation: Xianrui Lyu; supervision, funding acquisition and writing&#x2014;review and editing: Xiaodan Ren. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on request from the authors.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>
<p>To elucidate the physical significance of the SHAP-identified latent features, we conducted a systematic sensitivity analysis by perturbing individual dimensions while holding others constant. The procedure was as follows: a baseline latent vector was first randomly sampled from the latent space to serve as the origin. For each of the 16 SHAP-ranked dimensions, we generated a sequence of seven latent vectors by varying the target coordinate in increments of 0.5, spanning the range <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mtext>&#x0394;</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1.5</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>1.0</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mo>+</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mo>+</mml:mo><mml:mn>1.0</mml:mn><mml:mo>,</mml:mo><mml:mo>+</mml:mo><mml:mn>1.5</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, while keeping all other coordinates fixed at their baseline values. These perturbed latent vectors were subsequently decoded using the trained VAE to reconstruct the corresponding microstructural images, thereby enabling a direct visualization of how variation in each latent dimension influences microstructural morphology.</p>
<p>The results demonstrate a remarkable consistency with the SHAP-based sensitivity analysis. As illustrated in <xref ref-type="fig" rid="fig-16">Fig. A1a</xref>, an increase in the value of Feature 34 leads to a pronounced reduction in the volume fraction of circular inclusions, which also become more sparsely distributed. Conversely, a decrease in Feature 34 results in a substantial increase in inclusion volume fraction. It is worth noting that while variations in Feature 34 also induce other microstructural changes (e.g., inclusion arrangement), the alteration in volume fraction is the most visually salient effect. This visualization directly validates the negative correlation between Feature 34 identified by the SHAP analysis and homogenized stress.</p>
<fig id="fig-16">
<label>Figure A1</label>
<caption>
<title>Microstructural evolution induced by variations in latent features identified through SHAP analysis (<bold>a</bold>) Effect of Features 34 variation on microstructure; (<bold>b</bold>) Effect of Features 151 variation on microstructure; (<bold>c</bold>) Effect of Features 7 variation on microstructure; (<bold>d</bold>) Effect of Features 213 variation on microstructure; (<bold>e</bold>) Effect of Features 47 variation on microstructure; (<bold>f</bold>) Effect of Features 38 variation on microstructure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-16a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_74388-fig-16b.tif"/>
</fig>
<p>Similarly, as shown in <xref ref-type="fig" rid="fig-16">Fig. A1d</xref>, an increase in Feature 213 leads to a marked increase in the volume fraction of inclusions, which aligns perfectly with the positive correlation identified by the SHAP analysis. These findings confirm that the latent representations learned by the VAE encapsulate physically interpretable microstructural characteristics, and that the SHAP-identified dimensions correspond to meaningful geometric controls. The strong agreement between the perturbation-based visualizations and the SHAP dependency plots validates the reliability of the interpretability framework and underscores the utility of the latent space for guiding microstructure-sensitive design.</p>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Barile</surname> <given-names>C</given-names></string-name>, <string-name><surname>Casavola</surname> <given-names>C</given-names></string-name>, <string-name><surname>De Cillis</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Mechanical comparison of new composite materials for aerospace applications</article-title>. <source>Compos Part B Eng</source>. <year>2019</year>;<volume>162</volume>:<fpage>122</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compositesb.2018.10.101</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Le Duigou</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>3D printed continuous fiber reinforced composite lightweight structures: a review and outlook</article-title>. <source>Compos Part B Eng</source>. <year>2023</year>;<volume>250</volume>:<fpage>110450</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compositesb.2022.110450</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiao</surname> <given-names>JK</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>XY</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>JL</given-names></string-name>, <string-name><surname>Sheng</surname> <given-names>LY</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>YM</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>JH</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A review of research progress on machining carbon fiber-reinforced composites with lasers</article-title>. <source>Micromachines</source>. <year>2023</year>;<volume>14</volume>(<issue>1</issue>):<fpage>24</fpage>. doi:<pub-id pub-id-type="doi">10.3390/mi14010024</pub-id>; <pub-id pub-id-type="pmid">36677085</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Durandet</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Ruan</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Additively manufactured fiber-reinforced composites: a review of mechanical behavior and opportunities</article-title>. <source>J Mater Sci Technol</source>. <year>2022</year>;<volume>119</volume>:<fpage>219</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jmst.2021.11.063</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brockenbrough</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Suresh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wienecke</surname> <given-names>HA</given-names></string-name></person-group>. <article-title>Deformation of metal-matrix composites with continuous fibers: geometrical effects of fiber distribution and shape</article-title>. <source>Acta Metall Mater</source>. <year>1991</year>;<volume>39</volume>(<issue>5</issue>):<fpage>735</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0956-7151(91)90274-5</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Christman</surname> <given-names>T</given-names></string-name>, <string-name><surname>Needleman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Suresh</surname> <given-names>S</given-names></string-name></person-group>. <article-title>An experimental and numerical study of deformation in metal-ceramic composites</article-title>. <source>Acta Metall</source>. <year>1989</year>;<volume>37</volume>(<issue>11</issue>):<fpage>3029</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1016/0001-6160(89)90339-8</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Mishnaevsky</surname> <given-names>JL</given-names></string-name></person-group>. <source>Computational mesomechanics of composites</source>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons, Ltd.</publisher-name>; <year>2008</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Mechanical field guiding structure design strategy for meta-fiber reinforced hydrogel composites by deep learning</article-title>. <source>Adv Sci</source>. <year>2024</year>;<volume>11</volume>(<issue>22</issue>):<fpage>2310141</fpage>. doi:<pub-id pub-id-type="doi">10.1002/advs.202310141</pub-id>; <pub-id pub-id-type="pmid">38520708</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bostanabad</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kearney</surname> <given-names>T</given-names></string-name>, <string-name><surname>Brinson</surname> <given-names>LC</given-names></string-name>, <string-name><surname>Apley</surname> <given-names>DW</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Computational microstructure characterization and reconstruction: review of the state-of-the-art techniques</article-title>. <source>Prog Mater Sci</source>. <year>2018</year>;<volume>95</volume>(<issue>4</issue>):<fpage>1</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.pmatsci.2018.01.005</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Torquato</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Haslach</surname> <given-names>JHW</given-names></string-name></person-group>. <article-title>Random heterogeneous materials: microstructure and macroscopic properties</article-title>. <source>Appl Mech Rev</source>. <year>2002</year>;<volume>55</volume>(<issue>4</issue>):<fpage>B62</fpage>&#x2013;<lpage>3</lpage>. doi:<pub-id pub-id-type="doi">10.1115/1.1483342</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kenney</surname> <given-names>B</given-names></string-name>, <string-name><surname>Valdmanis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baker</surname> <given-names>C</given-names></string-name>, <string-name><surname>Pharoah</surname> <given-names>JG</given-names></string-name>, <string-name><surname>Karan</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Computation of TPB length, surface area and pore size from numerical reconstruction of composite solid oxide fuel cell electrodes</article-title>. <source>J Power Sources</source>. <year>2009</year>;<volume>189</volume>(<issue>2</issue>):<fpage>1051</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jpowsour.2008.12.145</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pattan</surname> <given-names>PC</given-names></string-name>, <string-name><surname>Mytri</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hiremath</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Classification of cast iron based on graphite grain morphology using neural network approach</article-title>. In: <conf-name>SPIE&#x2014;the international society for optical engineering</conf-name>. <publisher-loc>Washington, DC, USA</publisher-loc>: <publisher-name>SPIE</publisher-name>; <year>2010</year>. <fpage>75462S</fpage> p.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tewari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gokhale</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Nearest-neighbor distances between particles of finite size in three-dimensional uniform random microstructures</article-title>. <source>Mater Sci Eng A</source>. <year>2004</year>;<volume>385</volume>(<issue>1&#x2013;2</issue>):<fpage>332</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.msea.2004.06.049</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Brinson</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name></person-group>. <article-title>A descriptor-based design methodology for developing heterogeneous microstructural materials system</article-title>. <source>J Mech Des</source>. <year>2014</year>;<volume>136</volume>(<issue>5</issue>):<fpage>051007</fpage>. doi:<pub-id pub-id-type="doi">10.1115/1.4026649</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>IJ</given-names></string-name>, <string-name><surname>Pouget-Abadie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Warde-Farley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ozair</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Generative adversarial networks</article-title>. <comment>arXiv:1406.2661. 2014</comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational bayes</article-title>. <comment>arXiv:1312.6114. 2013</comment>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Improving direct physical properties prediction of heterogeneous materials from imaging data via convolutional neural network and a morphology-aware generative model</article-title>. <source>Comput Mater Sci</source>. <year>2018</year>;<volume>150</volume>(<issue>1</issue>):<fpage>212</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.commatsci.2018.03.074</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sardeshmukh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Reddy</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gautham</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bhattacharyya</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Material microstructure design using VAE-regression with multimodal prior</article-title>. <comment>arXiv:2402.17806. 2024</comment>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Seibert</surname> <given-names>P</given-names></string-name>, <string-name><surname>Otto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ra&#x00DF;loff</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ambati</surname> <given-names>M</given-names></string-name>, <string-name><surname>K&#x00E4;stner</surname> <given-names>M</given-names></string-name></person-group>. <article-title>DA-VEGAN: differentiably augmenting VAE-GAN for microstructure reconstruction from extremely small data sets</article-title>. <source>Comput Mater Sci</source>. <year>2024</year>;<volume>232</volume>(<issue>1</issue>):<fpage>112661</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.commatsci.2023.112661</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>YC</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Deep generative modeling for mechanistic-based learning and design of metamaterial systems</article-title>. <source>Comput Methods Appl Mech Eng</source>. <year>2020</year>;<volume>372</volume>:<fpage>113377</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cma.2020.113377</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lew</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Buehler</surname> <given-names>MJ</given-names></string-name></person-group>. <article-title>Encoding and exploring latent design space of optimal material structures via a VAE-LSTM model</article-title>. <source>Forces Mech</source>. <year>2021</year>;<volume>5</volume>:<fpage>100054</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.finmec.2021.100054</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kusampudi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Diehl</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Inverse design of dual-phase steel microstructures using generative machine learning model and Bayesian optimization</article-title>. <source>Int J Plast</source>. <year>2023</year>;<volume>171</volume>:<fpage>103776</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ijplas.2023.103776</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xue</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wallin</surname> <given-names>TJ</given-names></string-name>, <string-name><surname>Menguc</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Adriaenssens</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chiaramonte</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Machine learning generative models for automatic design of multi-material 3D printed composite solids</article-title>. <source>Extrem Mech Lett</source>. <year>2020</year>;<volume>41</volume>:<fpage>100992</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eml.2020.100992</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Guan</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Efficient property-oriented design of composite layups via controllable latent features using generative VAE</article-title>. <source>Compos Sci Technol</source>. <year>2025</year>;<volume>259</volume>:<fpage>110936</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compscitech.2024.110936</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>HG</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>DB</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Kwon</surname> <given-names>HY</given-names></string-name>, <string-name><surname>Won</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Optimization of physical quantities in the autoencoder latent space</article-title>. <source>Sci Rep</source>. <year>2022</year>;<volume>12</volume>(<issue>1</issue>):<fpage>9003</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-022-13007-5</pub-id>; <pub-id pub-id-type="pmid">35637207</pub-id></mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Prince</surname> <given-names>SJD</given-names></string-name></person-group>. <source>Understanding deep learning</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT Press</publisher-name>; <year>2023</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Machine learning-enabled constrained multi-objective design of architected materials</article-title>. <source>Nat Commun</source>. <year>2023</year>;<volume>14</volume>(<issue>1</issue>):<fpage>6630</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41467-023-42415-y</pub-id>; <pub-id pub-id-type="pmid">37857648</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lyu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Variational autoencoder guided conditional diffusion generative model for material microstructure reconstruction and inverse design</article-title>. <source>Mater Today Commun</source>. <year>2025</year>;<volume>48</volume>(<issue>1</issue>):<fpage>113087</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.mtcomm.2025.113087</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Michelucci</surname> <given-names>U</given-names></string-name></person-group>. <article-title>An introduction to autoencoders</article-title>. <comment>arXiv:2201.03898. 2022</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sohl-Dickstein</surname> <given-names>J</given-names></string-name>, <string-name><surname>Weiss</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Maheswaranathan</surname> <given-names>N</given-names></string-name>, <string-name><surname>Ganguli</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Deep unsupervised learning using nonequilibrium thermodynamics</article-title>. <comment>arXiv:1503.03585. 2015</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic models</article-title>. <comment>arXiv:2006.11239. 2020</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>D&#x00FC;reth</surname> <given-names>C</given-names></string-name>, <string-name><surname>Seibert</surname> <given-names>P</given-names></string-name>, <string-name><surname>R&#x00FC;cker</surname> <given-names>D</given-names></string-name>, <string-name><surname>Handford</surname> <given-names>S</given-names></string-name>, <string-name><surname>K&#x00E4;stner</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gude</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Conditional diffusion-based microstructure reconstruction</article-title>. <source>Mater Today Commun</source>. <year>2023</year>;<volume>35</volume>:<fpage>105608</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.mtcomm.2023.105608</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lyu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Microstructure reconstruction of 2D/3D random materials via diffusion-based deep generative models</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>(<issue>1</issue>):<fpage>5041</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-54861-9</pub-id>; <pub-id pub-id-type="pmid">38424207</pub-id></mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vlassis</surname> <given-names>NN</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Denoising diffusion algorithm for inverse design of microstructures with fine-tuned nonlinear material properties</article-title>. <source>Comput Methods Appl Mech Eng</source>. <year>2023</year>;<volume>413</volume>(<issue>1</issue>):<fpage>116126</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cma.2023.116126</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pandey</surname> <given-names>K</given-names></string-name>, <string-name><surname>Mukherjee</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rai</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>A</given-names></string-name></person-group>. <article-title>DiffuseVAE: efficient, controllable and high-fidelity generation from low-dimensional latents</article-title>. <comment>arXiv:2201.00308. 2022</comment>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Honarmandi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Attari</surname> <given-names>V</given-names></string-name>, <string-name><surname>Arroyave</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Accelerated materials design using batch Bayesian optimization: a case study for solving the inverse problem from materials microstructure to process specification</article-title>. <source>Comput Mater Sci</source>. <year>2022</year>;<volume>210</volume>(<issue>1</issue>):<fpage>111417</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.commatsci.2022.111417</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matsuda</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ookawara</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yasuda</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yoshikawa</surname> <given-names>S</given-names></string-name>, <string-name><surname>Matsumoto</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Framework for discovering porous materials: structural hybridization and Bayesian optimization of conditional generative adversarial network</article-title>. <source>Digit Chem Eng</source>. <year>2022</year>;<volume>5</volume>:<fpage>100058</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dche.2022.100058</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shahriari</surname> <given-names>B</given-names></string-name>, <string-name><surname>Swersky</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Adams</surname> <given-names>RP</given-names></string-name>, <string-name><surname>de Freitas</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Taking the human out of the loop: a review of bayesian optimization</article-title>. <source>Proc IEEE</source>. <year>2016</year>;<volume>104</volume>(<issue>1</issue>):<fpage>148</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jproc.2015.2494218</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tuo</surname> <given-names>R</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Inverse design of nonlinear mechanics of bio-inspired materials through interface engineering and Bayesian optimization</article-title>. <source>Extrem Mech Lett</source>. <year>2025</year>;<volume>78</volume>:<fpage>102359</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eml.2025.102359</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Altschuh</surname> <given-names>P</given-names></string-name>, <string-name><surname>Santoki</surname> <given-names>J</given-names></string-name>, <string-name><surname>Griem</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tosato</surname> <given-names>G</given-names></string-name>, <string-name><surname>Selzer</surname> <given-names>M</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Characterization of porous membranes using artificial neural networks</article-title>. <source>Acta Mater</source>. <year>2023</year>;<volume>253</volume>:<fpage>118922</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.actamat.2023.118922</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Schmitt</surname> <given-names>S</given-names></string-name>, <string-name><surname>Olhofer</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Recent advances in Bayesian optimization</article-title>. <comment>arXiv:2206.03301. 2022</comment>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rasmussen</surname> <given-names>CE</given-names></string-name></person-group>. <chapter-title>Gaussian processes in machine learning</chapter-title>. In: <source>Advanced lectures on machine learning</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2003</year>. p. <fpage>63</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-540-28650-9_4</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mo&#x0010D;kus</surname> <given-names>J</given-names></string-name></person-group>. <article-title>On Bayesian methods for seeking the extremum</article-title>. In: <conf-name>IFIP technical conference on optimization
techniques</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>1974</year>. p. <fpage>400</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xing</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Expected improvement for expensive optimization: a review</article-title>. <source>J Glob Optim</source>. <year>2020</year>;<volume>78</volume>(<issue>3</issue>):<fpage>507</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10898-020-00923-x</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lundberg</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>SI</given-names></string-name></person-group>. <article-title>A unified approach to interpreting model predictions</article-title>. <comment>arXiv:1705.07874. 2017</comment>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kondo</surname> <given-names>R</given-names></string-name>, <string-name><surname>Yamakawa</surname> <given-names>S</given-names></string-name>, <string-name><surname>Masuoka</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tajima</surname> <given-names>S</given-names></string-name>, <string-name><surname>Asahi</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Microstructure recognition using convolutional neural networks for prediction of ionic conductivity in ceramics</article-title>. <source>Acta Mater</source>. <year>2017</year>;<volume>141</volume>:<fpage>29</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.actamat.2017.09.004</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prifling</surname> <given-names>B</given-names></string-name>, <string-name><surname>R&#x00F6;ding</surname> <given-names>M</given-names></string-name>, <string-name><surname>Townsend</surname> <given-names>P</given-names></string-name>, <string-name><surname>Neumann</surname> <given-names>M</given-names></string-name>, <string-name><surname>Schmidt</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Large-scale statistical learning for mass transport prediction in porous materials using 90,000 artificially generated microstructures</article-title>. <source>Front Mater</source>. <year>2021</year>;<volume>8</volume>:<fpage>786502</fpage>. doi:<pub-id pub-id-type="doi">10.3389/fmats.2021.786502</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Denoising diffusion implicit models</article-title>. <comment>arXiv:2010.02502. 2020</comment>.</mixed-citation></ref>
</ref-list>
</back></article>














