<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">79254</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.079254</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>SWAGE-3D: Spectral Wasserstein Attention Generative Ensemble, A Comparative Analysis on the ShapeNet Dataset</article-title>
<alt-title alt-title-type="left-running-head">SWAGE-3D: Spectral Wasserstein Attention Generative Ensemble, A Comparative Analysis on the ShapeNet Dataset</alt-title>
<alt-title alt-title-type="right-running-head">SWAGE-3D: Spectral Wasserstein Attention Generative Ensemble, A Comparative Analysis on the ShapeNet Dataset</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Serin</surname><given-names>Zafer</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>zafer.serin@bilecik.edu.tr</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Karakuzu</surname><given-names>Cihan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Y&#x00FC;zge&#x00E7;</surname><given-names>U&#x011F;ur</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Pazaryeri Vocational School, Bilecik Seyh Edebali University</institution>, <addr-line>Bilecik</addr-line>, <country>T&#x00FC;rkiye</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Engineering, Bilecik Seyh Edebali University</institution>, <addr-line>Bilecik</addr-line>, <country>T&#x00FC;rkiye</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Zafer Serin. Email: <email>zafer.serin@bilecik.edu.tr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>5</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>2</issue>
<elocation-id>30</elocation-id>
<history>
<date date-type="received">
<day>18</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_79254.pdf"></self-uri>
<abstract>
<p>This study proposes SWAGE-3D (Spectral Wasserstein Attention Generative Ensemble), an enhanced 3D-VAE-GAN framework for single-view 3D object reconstruction using voxel-based representations. The proposed model integrates RGB-D encoding, Wasserstein adversarial learning with hybrid Lipschitz regularization, and a self-attention&#x2013;augmented generator to improve structural coherence and training stability. By combining variational latent modeling with stabilized Wasserstein optimization, the framework aims to address common challenges in 3D generative modeling, including mode collapse, unstable convergence, and insufficient global consistency. The encoder employs a depth-aware feature extraction strategy, while the discriminator utilizes a hybrid spectral normalization and gradient penalty mechanism to ensure robust approximation of the Wasserstein objective. Additionally, an ensemble strategy is applied at inference time to enhance reconstruction reliability. The proposed approach is evaluated on the ShapeNet dataset across 13 object categories using the Intersection over Union (IoU) metric. Experimental results demonstrate a 16.7% improvement over the baseline 3D-VAE-GAN and competitive performance against state-of-the-art voxel-based reconstruction methods. These findings confirm that the synergistic integration of depth cues, stabilized Wasserstein training, and attention mechanisms significantly enhances single-view 3D reconstruction performance.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Three-dimensional reconstruction</kwd>
<kwd>variational autoencoder</kwd>
<kwd>generative adversarial network</kwd>
<kwd>depth estimation</kwd>
<kwd>residual neural network</kwd>
<kwd>ensemble learning</kwd>
<kwd>attention mechanism</kwd>
</kwd-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Single-view 3D object reconstruction remains a fundamental challenge in computer vision due to the inherently ill-posed nature of inferring volumetric geometry from a single 2D observation. The problem has broad applications in virtual reality, robotics, and digital content creation. Traditional manual 3D modeling is time-consuming and often requires substantial expert effort for producing a single high-quality model [<xref ref-type="bibr" rid="ref-1">1</xref>]. Consequently, learning-based reconstruction methods have emerged as scalable alternatives for accelerating 3D content generation.</p>
<p>Among various 3D representations, voxel-based models provide a structured volumetric formulation that integrates naturally with convolutional neural networks. As illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, voxel grids encode occupancy information within a regular 3D lattice, enabling direct application of 3D convolutions and probabilistic generative modeling [<xref ref-type="bibr" rid="ref-2">2</xref>]. In contrast, point-cloud representations lack explicit surface connectivity and often require post-processing for surface reconstruction [<xref ref-type="bibr" rid="ref-3">3</xref>], while polygon mesh representations demand predefined topology and may suffer from discretization artifacts [<xref ref-type="bibr" rid="ref-4">4</xref>]. Due to their structural regularity and compatibility with convolutional architectures, voxel representations remain a practical choice for end-to-end reconstruction pipelines.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Voxel, point-cloud, and polygon (mesh) representations of a sphere.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-1.tif"/>
</fig>
<p>Deep learning has substantially advanced image-based feature extraction and generative modeling. Convolutional Neural Networks (CNNs), including architectures such as ResNet [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>], have demonstrated strong capability in visual representation learning. The inclusion of diverse training strategies and augmentation techniques, such as MixUp, has further improved generalization performance [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Generative frameworks, particularly Variational Autoencoders (VAEs) [<xref ref-type="bibr" rid="ref-11">11</xref>] and Generative Adversarial Networks (GANs) [<xref ref-type="bibr" rid="ref-12">12</xref>], have enabled probabilistic modeling of volumetric data. However, standard GAN training is prone to instability and mode collapse, especially in high-dimensional voxel spaces [<xref ref-type="bibr" rid="ref-13">13</xref>]. Wasserstein GAN formulations with Gradient Penalty (WGAN-GP) have been proposed to improve gradient behavior and convergence stability [<xref ref-type="bibr" rid="ref-14">14</xref>]. Moreover, attention mechanisms have demonstrated the ability to model long-range dependencies and improve structural coherence in generative tasks [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>More recently, diffusion-based generative models have significantly advanced single-view 3D reconstruction. Approaches such as Zero-1-to-3 [<xref ref-type="bibr" rid="ref-17">17</xref>] and DreamFusion [<xref ref-type="bibr" rid="ref-18">18</xref>] leverage diffusion priors to synthesize novel views or optimize neural radiance fields, enabling high-fidelity implicit 3D representations. Similarly, Shap-E [<xref ref-type="bibr" rid="ref-19">19</xref>] and Point-E [<xref ref-type="bibr" rid="ref-20">20</xref>] demonstrate the effectiveness of diffusion-guided implicit and point-based generation frameworks. While these approaches achieve impressive visual realism, they typically rely on implicit representations or multi-stage optimization pipelines, which differ from voxel-based end-to-end adversarial reconstruction frameworks.</p>
<p>Beyond architectural components, effective optimization strategies play a critical role in volumetric generation. Adaptive learning rate scheduling techniques [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>], ensemble learning approaches [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>], and computational optimization methods such as automatic mixed precision have been shown to improve efficiency and training robustness in deep neural networks [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<p>Despite these advances, several challenges persist in single-view voxel reconstruction: (i) limited geometric cues in RGB-only inputs, (ii) adversarial training instability in high-dimensional volumetric generation, and (iii) variance in reconstruction quality across training epochs. Addressing these issues requires a carefully stabilized generative framework that integrates complementary strategies within a unified reconstruction pipeline.</p>
<p>To this end, we propose SWAGE-3D (Spectral Wasserstein Attention Generative Ensemble for 3D Modeling), a stabilized RGB-D VAE-GAN framework designed for single-view voxel reconstruction. The proposed approach integrates depth-assisted feature encoding, Wasserstein adversarial training with spectral regularization, attention-based structural modeling, and checkpoint-based ensemble inference. Rather than introducing entirely new building blocks, SWAGE-3D systematically integrates and validates complementary stabilization strategies to improve reconstruction consistency and volumetric coherence.</p>
<p><bold><italic>Contributions</italic></bold></p>
<p>The main contributions of this work are summarized as follows:<list list-type="bullet">
<list-item>
<p>Depth-Assisted RGB-D Encoding: A four-channel RGB-D input strategy is introduced by integrating monocular depth maps into the reconstruction pipeline. A ResNet18 encoder is adapted to accommodate depth information via mean-initialized channel expansion, enabling stable transfer learning.</p></list-item>
<list-item>
<p>Stabilized Adversarial Volumetric Training: A Wasserstein GAN with Gradient Penalty is combined with spectral normalization to improve convergence stability and control Lipschitz continuity in high-dimensional voxel generation.</p></list-item>
<list-item>
<p>Attention-Augmented Generator: A self-attention mechanism is incorporated to model long-range spatial dependencies, enhancing structural coherence in reconstructed voxel grids.</p></list-item>
<list-item>
<p>Checkpoint-Based Ensemble Inference: An IoU-weighted ensemble strategy is employed to reduce prediction variance and improve inference robustness.</p></list-item>
<list-item>
<p>Comprehensive Experimental Validation: Extensive ablation studies, sensitivity analysis, and multi-metric evaluations are conducted to quantify the individual and collective impact of the proposed components.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Recent advancements in 3D reconstruction have been shaped by diverse representation paradigms and learning strategies, ranging from voxel-based modeling to implicit function learning and depth-assisted inference mechanisms. This section organizes prior studies according to representation paradigms and methodological focus.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Voxel-Based and Generative 3D Reconstruction</title>
<p>Voxel representations have long been adopted for structured 3D modeling due to their regular grid formulation and compatibility with convolutional neural networks. Early works explored visual similarity and rotation-invariant normalization for 3D retrieval [<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>The emergence of deep generative modeling significantly advanced volumetric synthesis. Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established probabilistic and adversarial paradigms for high-dimensional data generation. Building upon these foundations, Wu et al. introduced 3D-GAN and 3D-VAE-GAN [<xref ref-type="bibr" rid="ref-28">28</xref>], enabling volumetric generation from both latent codes and 2D images. Recent advancements have sought to unify these paradigms; for instance, UniRecGen [<xref ref-type="bibr" rid="ref-29">29</xref>] integrates feed-forward reconstruction with diffusion-based generation to achieve highly consistent multi-view 3D models.</p>
<p>Recurrent aggregation mechanisms were later proposed for multi-view reconstruction in 3D-R2N2 [<xref ref-type="bibr" rid="ref-30">30</xref>]. Improvements in adversarial stability and scalable generative modeling were introduced through Wasserstein formulations like 3D-IWGAN [<xref ref-type="bibr" rid="ref-31">31</xref>] and geometry-aware 3D GANs [<xref ref-type="bibr" rid="ref-32">32</xref>]. Additional voxel-based generative and classification studies provided further architectural insights into 3D representation learning and recognition [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-37">37</xref>]. In particular, orientation-aware modeling strategies also enriched the design space of 3D deep learning frameworks [<xref ref-type="bibr" rid="ref-38">38</xref>].</p>
<p>More recent voxel-based frameworks focus on improved feature fusion and long-range dependency modeling. Pix2Vox [<xref ref-type="bibr" rid="ref-39">39</xref>] introduced context-aware fusion for combining coarse reconstructions, while TMVNet [<xref ref-type="bibr" rid="ref-40">40</xref>] leveraged Transformer-based encoders for enhanced multi-view feature aggregation. Similarly, VoxFormer [<xref ref-type="bibr" rid="ref-41">41</xref>] demonstrated the power of sparse voxel transformers for camera-based volumetric scene completion. To further refine these architectures, SS3DNet-AF [<xref ref-type="bibr" rid="ref-42">42</xref>] proposed an attention-based fusion mechanism that enhances the network&#x2019;s focus on relevant geometric details from a single view. Similarly, recent efforts have directly upgraded multi-view pipelines by integrating multi-head attention refiners to reduce boundary prediction errors and capture intricate structural nuances [<xref ref-type="bibr" rid="ref-43">43</xref>]. Attention-enhanced voxel models such as SV3D-CDFF [<xref ref-type="bibr" rid="ref-44">44</xref>] and Semantic Voxel Structure (SVS) [<xref ref-type="bibr" rid="ref-45">45</xref>] further addressed discrepancies between image and voxel domains.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Alternative Representations: Point-Based and Implicit Methods</title>
<p>To overcome voxel resolution limitations, alternative 3D representations have been explored. Point-cloud-based generative models employed Earth Mover&#x2019;s Distance and deep autoencoding strategies for shape generation and completion [<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]. DescriptorNet introduced an energy-based probabilistic model for 3D pattern synthesis [<xref ref-type="bibr" rid="ref-48">48</xref>].</p>
<p>Implicit representations further advanced continuous shape modeling. Occupancy Networks [<xref ref-type="bibr" rid="ref-49">49</xref>] represented 3D objects as continuous decision functions, enabling high-resolution reconstruction without explicit voxel discretization. IF-NET [<xref ref-type="bibr" rid="ref-50">50</xref>] extended implicit learning for shape completion. While these methods, alongside recent accelerated rendering advancements [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>], alleviate discretization constraints, voxel representations remain advantageous for direct convolutional processing and adversarial volumetric learning.</p>
<p>Beyond implicit and point-based paradigms, studies have explored deformable surface modeling and graph-based generative approaches for structured 3D synthesis. Multi-chart surface parameterization and surface-oriented generation methods demonstrated that complex 3D geometry can be reconstructed through mesh- or chart-based learning strategies [<xref ref-type="bibr" rid="ref-53">53</xref>&#x2013;<xref ref-type="bibr" rid="ref-55">55</xref>]. Studies on dynamic reconstruction, semantic shape modeling, and volumetric autoencoding supported the feasibility of learning structured geometric representations under diverse architectural assumptions [<xref ref-type="bibr" rid="ref-56">56</xref>&#x2013;<xref ref-type="bibr" rid="ref-59">59</xref>]. While these approaches offer advantages in continuous surface modeling, they often require predefined topology constraints or specialized deformation priors. In contrast, voxel-based representations provide a regular volumetric grid structure that facilitates stable adversarial training and unified convolutional processing within a single-view reconstruction pipeline.</p>
<p>Recent advances in point-based and geometric learning have further expanded 3D representation capabilities. Transformer-based architectures and Vision Transformer (ViT) adaptations have been applied to point cloud processing and large-scale generalizable reconstruction to model global spatial dependencies more effectively [<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>]. In parallel, geometric neural operator frameworks have emerged as powerful tools for learning structured mappings in high-dimensional geometric domains [<xref ref-type="bibr" rid="ref-62">62</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. These developments provide promising alternatives for geometric representation learning, complementing voxel-based approaches in scenarios where continuous or sparse representations are preferred. Transformer-based 3D reconstruction has also been extended toward neural implicit surface modeling. For example, SparseNeuS [<xref ref-type="bibr" rid="ref-64">64</xref>] integrates transformer architectures with neural implicit surfaces to enhance geometric consistency across sparse views, demonstrating improved structural coherence in continuous representations.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Depth Integration and Stabilization Strategies</title>
<p>Depth information has increasingly been incorporated to enhance geometric inference. Multi-view supervision methods such as Differentiable Ray Consistency (DRC) [<xref ref-type="bibr" rid="ref-65">65</xref>] demonstrated that volumetric predictions can be learned without explicit 3D labels. Recent advances in monocular depth estimation, particularly Depth Anything V2 [<xref ref-type="bibr" rid="ref-66">66</xref>], have significantly improved robustness and generalization across diverse environments.</p>
<p>Architectural refinements have also focused on stability and feature modeling. Self-attention mechanisms and Transformer-based encoders have been introduced to better capture long-range spatial dependencies in 3D reconstruction [<xref ref-type="bibr" rid="ref-67">67</xref>&#x2013;<xref ref-type="bibr" rid="ref-69">69</xref>]. Ensemble learning strategies have been widely adopted to reduce prediction variance and enhance robustness in deep neural networks.</p>
<p>Despite substantial progress, several challenges persist in single-view voxel reconstruction, including limited geometric cues in RGB-only inputs, adversarial instability in high-dimensional volumetric generation, and reconstruction variance across training epochs. Although prior studies have addressed these issues individually, their combined and systematically stabilized integration within a unified RGB-D voxel-based generative framework has received comparatively limited investigation. In this context, the present study emphasizes the controlled integration and empirical validation of complementary stabilization strategies for depth-assisted voxel reconstruction.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Spectral Wasserstein Attention Generative Ensemble for 3D Model (SWAGE-3D)</title>
<sec id="s3_1">
<label>3.1</label>
<title>Overview of SWAGE-3D</title>
<p>Unlike our previous VAE-based approach [<xref ref-type="bibr" rid="ref-70">70</xref>], which primarily focused on transfer learning for latent representation stabilization, the proposed SWAGE-3D framework systematically integrates adversarial Wasserstein training, spectral normalization, self-attention modules, and ensemble learning to address stability and geometric coherence in single-view voxel reconstruction. SWAGE-3D framework builds upon the foundational work of Wu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] on 3D-VAE-GAN, introducing several key innovations to enhance both training stability and the quality of generated 3D models. The overall architecture, illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, consists of three primary network components that operate in a unified voxel-based generative pipeline.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Overall architecture of the proposed SWAGE-3D framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-2.tif"/>
</fig>
<p>The architecture integrates three primary network components that work together to deliver superior performance. The Encoder Network serves as the initial processing unit, taking 2D input images that include both RGB and depth information and mapping them into a latent space characterized by mean (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula>) and variance (<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula>) parameters. By incorporating depth maps alongside RGB data, the encoder captures richer geometric details, enabling more accurate and informative latent representations. The Generator Network then transforms latent vectors sampled from this space into detailed 3D voxel models using transposed 3D convolutions and a novel self-attention mechanism. This mechanism enhances the model&#x2019;s ability to capture long-range structural coherence. Finally, the Discriminator Network evaluates the generated 3D models to differentiate between real scanned objects and synthetic samples. By providing precise adversarial feedback based on the Wasserstein distance, this component guides the generator in producing outputs that are increasingly indistinguishable from real-world data. The algorithm for the proposed model is shown in Algorithm 1.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>RGB-D Encoder</title>
<p>The SWAGE-3D framework employs a depth-assisted RGB-D encoding strategy to enrich geometric feature representation from single-view inputs. Instead of relying solely on RGB images, monocular depth maps are generated and fused as an additional channel, forming a four-channel input representation. Depth maps are extracted from the input RGB images using the Depth Anything V2 model, which provides robust monocular depth estimation across diverse scenes. Each RGB image of size <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> is augmented with its corresponding depth map to construct a <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula> RGB-D tensor. This integration enables the encoder to capture complementary geometric cues that are not explicitly encoded in color information alone.</p>
<p>All inputs are normalized using ImageNet statistics for the RGB channels, while the depth channel is normalized using its dataset-wide mean and standard deviation to maintain scale consistency. A ResNet18 backbone is adopted as the encoder component. Since the standard ResNet18 architecture is pre-trained on3-channel RGB images, its first convolutional layer is modified to accept 4-channel RGB-D inputs. To preserve transfer learning benefits, the pre-trained weights corresponding to the three RGB channels are retained. The weights of the additional depth channel are initialized as the mean of the RGB channel weights: <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mtext>depth</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>3</mml:mn></mml:mfrac><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>R</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This initialization allows the depth channel to start from a balanced representation without disrupting learned feature distributions. The modified encoder outputs mean <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03BC;</mml:mi></mml:math></inline-formula> and log-variance <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula> parameters, from which latent vectors are sampled via the reparameterization trick:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>&#x03D5;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2299;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03B5;</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> stands for the random noise drawn from the standard normal distribution. This stochastic encoding enables variational learning while preserving stable gradient propagation during training.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Generative Backbone and Adversarial Stabilization</title>
<p>The generative backbone consists of a 3D convolutional generator and a Wasserstein-based discriminator. The generator transforms the sampled latent vector into a <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> voxel grid. To address instability commonly observed in standard GAN training, a Wasserstein GAN with Gradient Penalty (WGAN-GP) formulation is adopted [<xref ref-type="bibr" rid="ref-71">71</xref>]. The discriminator estimates the Wasserstein distance between real and generated voxel distributions. Its objective is defined as:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>D</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2207;</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:math></disp-formula></p>
<p>Although the Wasserstein distance formulation based on the Kantorovich&#x2013;Rubinstein (KR) duality provides a theoretically sound objective, its practical implementation remains an approximation [<xref ref-type="bibr" rid="ref-72">72</xref>]. In theory, the KR dual formulation considers the supremum over all 1-Lipschitz functions. However, the critic is restricted to a parameterized neural network with finite capacity. The learned critic provides only an approximation of the true Wasserstein-1 distance rather than its exact computation [<xref ref-type="bibr" rid="ref-72">72</xref>].</p>
<p>To mitigate these limitations, we adopt a hybrid stabilization strategy combining spectral normalization (global Lipschitz control) with a reduced-weight gradient penalty term (local gradient regularization). This dual constraint balances critic expressiveness and stability, improving approximation smoothness without severely restricting model capacity. This approach serves as a stabilized practical realization of Wasserstein-based adversarial learning. In our generator network, a self-attention mechanism is incorporated after the third layer to model long-range spatial dependencies, ensuring global consistency in the generated volumes.</p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-9.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Self-Attention Module</title>
<p>To enhance structural coherence in volumetric generation, a self-attention mechanism is incorporated into the generator. While convolutional layers effectively capture local spatial patterns, their receptive fields remain limited. In voxel-based reconstruction tasks, long-range spatial dependencies are critical for maintaining global object consistency. The self-attention module enables direct interactions between distant spatial regions within the volumetric feature maps. For an intermediate 3D feature tensor <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>F</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, query, key, and value representations are obtained via <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolutions. The attention map is computed based on query-key similarity, and the resulting representation is combined with the original feature tensor through a residual connection: <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mtext>Attention</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>F</mml:mi></mml:math></inline-formula>. Here, <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is a learnable scalar parameter initialized to zero to stabilize early training.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Ensemble Strategy</title>
<p>To enhance inference robustness and mitigate performance fluctuations across training epochs, a checkpoint-based ensemble strategy is employed during the testing phase. Instead of relying on a single trained model, multiple late-stage checkpoints are aggregated to produce a more stable and accurate voxel reconstruction. Model performance was monitored using the IoU metric, and a stabilization plateau was observed between epochs 125 and 149, indicating convergence with minor oscillations. From this stable region, epochs 129, 139, and 149 were selected to construct the ensemble.</p>
<p>Let <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the voxel prediction produced by the <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>i</mml:mi></mml:math></inline-formula>-th checkpoint model, and let <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> represent its validation IoU score. The final ensemble prediction is computed as a weighted aggregation: <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mrow><mml:mover><mml:mi>V</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>G</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2211;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>. Following aggregation, a voxel occupancy threshold <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> is applied to obtain the final binary volumetric reconstruction. This ensemble mechanism reduces variance introduced by stochastic training dynamics and improves overall reconstruction accuracy.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Composite Loss Function and Training Dynamics</title>
<p>The optimization of SWAGE-3D is formulated as a multi-stage composite objective that jointly regulates latent space regularity, volumetric reconstruction fidelity, and adversarial distribution alignment. The framework simultaneously optimizes three interacting networks: the encoder (<italic>E</italic>), the generator (<italic>G</italic>), and the discriminator (<italic>D</italic>), using a coordinated update strategy that bridges Variational Autoencoder (VAE) and Generative Adversarial Network (GAN) paradigms.</p>
<p>The primary objective of the VAE stream is to ensure that the encoder-generator pair can effectively map a single-view RGB-D input to its corresponding 3D geometry. To ensure voxel-wise consistency between predicted grids and ground-truth occupancy, a Mean Squared Error (MSE) reconstruction loss is employed:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the predicted and ground-truth voxel grids, respectively. Simultaneously, the encoder is regularized using the Kullback&#x2013;Leibler (KL) divergence to enforce the latent distribution to follow a unit Gaussian prior <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>z</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-73">73</xref>]:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x2211;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>&#x03BC;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>This dual objective allows the encoder to capture a robust, continuous representation of the 3D shapes while the generator learns the fundamental volumetric mapping.</p>
<p>To enhance the realism and structural sharpness of the generated voxels beyond simple point-wise matching, an adversarial stream is integrated. Under the Wasserstein formulation, the discriminator acts as a critic that estimates the Earth-Mover distance between the generated distribution <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>P</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and the real data distribution <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>P</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula>. Its objective is defined as:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>D</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represents the gradient penalty term enforcing the 1-Lipschitz continuity constraint [<xref ref-type="bibr" rid="ref-74">74</xref>]. In our implementation, spectral normalization is applied to the critic&#x2019;s layers to provide global Lipschitz control, allowing the gradient penalty coefficient <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to be reduced to <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mn>1.0</mml:mn></mml:math></inline-formula>. This prevents the vanishing gradient problem and stabilizes the training of the generator.</p>
<p>The technical novelty of the SWAGE-3D training procedure lies in the coordinated update of the generator. The generator is not merely trained to fool the discriminator but is explicitly forced to satisfy both the VAE reconstruction constraint and the GAN adversarial requirement. The total objectives for the encoder (<inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>E</mml:mi></mml:msub></mml:math></inline-formula>) and generator (<inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula>) are defined as:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>E</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>G</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>By minimizing <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula>, the generator learns to produce voxels that are geometrically accurate relative to the input image (via <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) while simultaneously conforming to the global distribution of real 3D objects (via the adversarial term). In all experiments, the weights are set to <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>. To maintain equilibrium in this complex adversarial landscape, the discriminator is updated <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> times for every single update of the encoder and generator, ensuring a reliable approximation of the Wasserstein distance before each generative refinement step.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Setup</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Dataset</title>
<p>This study utilizes the ShapeNet dataset, a widely recognized and comprehensive repository of 3D models across various object categories. ShapeNet contains a total of 55 different object classes, making it one of the most extensively used datasets in the literature for 3D object generation and reconstruction tasks [<xref ref-type="bibr" rid="ref-75">75</xref>]. In alignment with prior studies, this research focuses on a subset of 13 objects selected from the ShapeNet dataset. These objects were chosen to ensure compatibility with existing benchmarks and facilitate meaningful comparisons.</p>
<p>The ShapeNet dataset provides 3D model files for each object, which can be voxelized using the Binvox method. The voxel term refers to a volumetric pixel, representing a point in the 3D environment where a cube is either present (state 1) or absent (state 0). Using tools developed by the ShapeNet team and other researchers, 2D images of the models can be generated from multiple viewpoints as desired. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> demonstrates an example of a 3D voxel representation alongside its corresponding 2D image for each object used in this study. The dataset is divided into 80% for training and 20% for testing for each object, ensuring a balanced evaluation framework.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>2D images and corresponding 3D models of 13 objects selected from the ShapeNet dataset.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-3.tif"/>
</fig>
<p>To ensure compatibility with the ResNet18 architecture, 2D images were resized to dimensions of 224 <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224 <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3. Here, the first two dimensions (224 <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224) represent the spatial resolution (horizontal and vertical pixel count), while the third dimension corresponds to the RGB channels (Red, Green, Blue). Each channel is represented by 8 bits, resulting in a total of 24 bits per image. While some datasets include an additional alpha channel (RGBA) to represent transparency, this study focuses exclusively on RGB images.</p>
<p>For 3D models, a voxelized representation is employed. Each model is expressed within a 32 <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 32 <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 32 grid (resolution), where cubes are either added or omitted at specific points in the 3D space. This binary representation ensures efficient processing while preserving the structural integrity of the objects.</p>
<p>Following preprocessing procedures, 2D input images with dimensions of 224 <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224 <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 3 and their corresponding 3D voxel objects of size 32 <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 32 <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 32 were prepared. Depth maps were extracted from the 2D images using the Depth Anything V2 architecture [<xref ref-type="bibr" rid="ref-66">66</xref>] and integrated as a fourth channel into the input. These depth maps are represented as 8-bit grayscale images, expanding the input dimensions to 224 <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 224 <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 4.</p>
<p>The incorporation of depth maps was motivated by several potential improvements. Specifically, this enhancement aims to provide the architecture with additional geometric information, spatial depth cues, and improved object boundary delineation. Furthermore, it mitigates challenges related to illumination variations, enriches feature representation, and strengthens the discriminative capabilities of the model. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates example inputs along with their corresponding depth maps extracted using the Depth Anything V2 architecture.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>2D images and corresponding depth maps using Depth Anything V2.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-4.tif"/>
</fig>
<p>During the integration of depth maps, specific normalization parameters were employed. The mean (<inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mn>0.330</mml:mn></mml:math></inline-formula>) and standard deviation (<inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mn>0.236</mml:mn></mml:math></inline-formula>) values used for depth channel normalization were empirically calculated across the entire training subset of the ShapeNet dataset to ensure optimal distribution for the ResNet18 encoder.</p>
<p>All experiments were conducted on the ShapeNetCore dataset using the 13 object categories described in this section. The dataset was divided into 80% training and 20% testing samples for each category to ensure balanced evaluation across object types. The 3D models were voxelized at a resolution of <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>, while the corresponding 2D input images were resized to <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula>. Monocular depth maps were generated for each RGB image and concatenated as a fourth channel, forming an RGB-D input representation.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Implementation Details</title>
<p>The SWAGE-3D framework was implemented using the PyTorch deep learning library. All experiments were conducted on a workstation equipped with an Intel i7-4790 CPU, 16 GB RAM, and a single NVIDIA GTX 980 Ti GPU.</p>
<p>To ensure reproducibility and training stability, we adopted a specific weight initialization and data augmentation strategy. The ResNet18 encoder backbone, pre-trained on ImageNet, was modified to accept four-channel RGB-D inputs. The weights for the additional depth channel were initialized as the mean of the pre-trained RGB weights, <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mtext>depth</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtext>avg</mml:mtext><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>R</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>G</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, to preserve the benefits of transfer learning. During training, 2D MixUp data augmentation was applied to the input images with an interpolation probability of <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> and a Beta distribution parameter of <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.2</mml:mn></mml:math></inline-formula> to mitigate mode collapse.</p>
<p>The final hyperparameters, determined through empirical validation and sensitivity analysis, are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>. All models were trained for 150 epochs with a batch size of 32. The latent space dimensionality was strictly set to <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>200</mml:mn></mml:math></inline-formula> to balance representational capacity and training stability.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Final training configuration and hyperparameters.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th>Parameter</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="4">Optimizer (Adam)</td>
<td>Generator Learning Rate (<inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>G</mml:mi></mml:msub></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mn>2.5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Discriminator Learning Rate (<inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>D</mml:mi></mml:msub></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mn>1.0</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Encoder Learning Rate (<inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>E</mml:mi></mml:msub></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mn>1.0</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Betas (<inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mo stretchy="false">(</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>0.9</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="4">Loss Weights</td>
<td>Reconstruction Weight (<inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</td>
<td>1.0</td>
</tr>
<tr>
<td>KL Divergence Weight (<inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</td>
<td>1.0</td>
</tr>
<tr>
<td>Gradient Penalty Weight (<inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>G</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</td>
<td>1.0&#x002A;</td>
</tr>
<tr>
<td>Critic Iterations (<inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>r</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</td>
<td>5</td>
</tr>
<tr>
<td rowspan="4">Stability</td>
<td>Latent Dimension (<inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>d</mml:mi><mml:mi>z</mml:mi></mml:msub></mml:math></inline-formula>)</td>
<td>200</td>
</tr>
<tr>
<td>Classifier Threshold (<inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>)</td>
<td>0.8</td>
</tr>
<tr>
<td>MixUp Prob/Alpha (<inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>p</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>)</td>
<td>0.5/0.2</td>
</tr>
<tr>
<td>Occupancy Threshold (<inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>)</td>
<td>0.5</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-1fn1" fn-type="other">
<p>Note: &#x002A;Reduced from 10.0 to 1.0 when Spectral Normalization is enabled.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Optimization was performed using the Adam optimizer. To prevent the discriminator from overpowering the generator, a balancing threshold was implemented: the discriminator was only updated if its classification accuracy was below <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mtext>thresh</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.8</mml:mn></mml:math></inline-formula>. Following the WGAN-GP strategy, the critic was updated five times per generator update (<inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mtext>critic</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>).</p>
<p>The composite loss function integrates voxel reconstruction loss (MSE), KL divergence (<inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula>), and adversarial feedback. In the baseline WGAN-GP configuration, <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>GP</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> was set to 10, but was reduced to 1.0 when spectral normalization was enabled to prevent over-constraining the critic. Model performance was monitored via Intersection over Union (IoU). As shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, stabilization occurs after epoch 125, motivating the selection of late-stage checkpoints (epochs 129, 139, and 149) for ensemble construction. For quantitative evaluation, probabilistic voxel outputs were binarized using a fixed occupancy threshold of <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Results and potential plateau regions for 150 epochs of training IoU for the objects Bench, Cabinet, Car, and Chair.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-5.tif"/>
</fig>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Sensitivity Analysis of Loss Weights</title>
<p>To evaluate the robustness of the SWAGE-3D framework, we conducted a sensitivity study on the primary loss weights, <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, using the Loudspeaker category as a representative sample. We tested scaling factors of <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mo>&#x00B1;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> relative to the default configuration. Due to computational constraints, each configuration was trained for 30 epochs to observe early convergence trends and training stability. The results are summarized in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Sensitivity analysis of loss weights on Loudspeaker category (evaluated at epoch 30).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th><inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th>Mean IoU</th>
<th>Stability</th>
</tr>
</thead>
<tbody>
<tr>
<td>Default (Optimal)</td>
<td>1.0</td>
<td>1.0</td>
<td>0.6676</td>
<td><bold>Stable</bold></td>
</tr>
<tr>
<td>Low Reconstruction</td>
<td>0.5</td>
<td>1.0</td>
<td>0.6617</td>
<td>Stable</td>
</tr>
<tr>
<td>High Reconstruction</td>
<td>2.0</td>
<td>1.0</td>
<td><bold>0.6766</bold></td>
<td>Minor Oscillations</td>
</tr>
<tr>
<td>Low KL Regularization</td>
<td>1.0</td>
<td>0.5</td>
<td>0.6600</td>
<td>High Variance</td>
</tr>
<tr>
<td>High KL Regularization</td>
<td>1.0</td>
<td>2.0</td>
<td>0.6695</td>
<td>Stable (Blurred)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-2fn1" fn-type="other">
<p>Note: Bold values indicate the best performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The results demonstrate that SWAGE-3D exhibits remarkable stability across different objective weightings, with IoU scores remaining within a narrow margin (<inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mo>&#x003C;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>2</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>). While the &#x2018;High Reconstruction&#x2019; (<inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula>) configuration yielded a slightly higher IoU of 0.6766 at this stage, it introduced minor oscillations in the discriminator loss, which could potentially destabilize the adversarial balance during full-scale 150-epoch training. Conversely, reducing <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to 0.5 led to increased variance in the Wasserstein critic&#x2019;s feedback, confirming that adequate latent regularization is essential for smooth adversarial mapping. Doubling the KL weight (<inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2.0</mml:mn></mml:math></inline-formula>) maintained stability but resulted in slightly smoother, less detailed voxel grids (posterior collapse). Based on these observations, the <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula> configuration was selected as the optimal balance for long-term training, ensuring both high geometric fidelity and consistent adversarial equilibrium across all object categories.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metrics</title>
<p>To quantitatively evaluate the reconstruction performance of SWAGE-3D and the compared methods, multiple complementary metrics were employed to assess volumetric overlap, structural accuracy, and geometric distance consistency. The primary evaluation metric is the Intersection over Union (IoU), which measures the volumetric overlap between predicted and ground-truth voxel grids. IoU is defined as:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msup><mml:mi>V</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the predicted and ground-truth voxel occupancy grids, respectively. This metric provides a rigorous assessment of spatial consistency between reconstructed and reference 3D shapes. For all evaluations, a voxel occupancy threshold of <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula> was applied to binarize the generator outputs.</p>
<p>While IoU evaluates total volumetric overlap, it may not fully capture the precision of reconstructed surface details. Therefore, the F-score (<inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>) is additionally reported to provide a balanced measure of reconstruction fidelity by calculating the harmonic mean of precision and recall based on voxel occupancy. The F-score is computed as:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>In this context, precision represents the accuracy of the predicted occupied voxels, while recall indicates the fraction of the ground-truth structure successfully recovered by the model. This metric is particularly effective in evaluating the reconstruction quality of thin and complex structures where volumetric overlap alone might be insufficient to reflect structural integrity.</p>
<p>To further quantify the geometric consistency between reconstructed and ground-truth shapes, the Chamfer Distance (<italic>CD</italic>) is employed. Given two point sets <italic>P</italic> and <italic>G</italic> derived from the coordinates of occupied voxels in the predicted and ground-truth grids, respectively, <italic>CD</italic> is defined as:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>C</mml:mi><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>G</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:munder><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>g</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>G</mml:mi></mml:mrow></mml:munder><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>g</mml:mi><mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>G</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>g</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>G</mml:mi></mml:mrow></mml:munder><mml:munder><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:munder><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>g</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>p</mml:mi><mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup></mml:math></disp-formula></p>
<p>Chamfer Distance evaluates bidirectional nearest-neighbor consistency and captures geometric deviations that influence surface accuracy even when they do not significantly affect the overall voxel overlap. By jointly reporting IoU, F-score, and Chamfer Distance, a comprehensive assessment of volumetric reconstruction accuracy, structural precision, and geometric consistency is provided.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Quantitative Comparison with Existing Methods</title>
<p>To evaluate its performance, SWAGE-3D is compared with several representative 3D reconstruction methods, including 3D-R2N2 [<xref ref-type="bibr" rid="ref-30">30</xref>], which employs Recurrent Neural Networks (RNNs) for single- or multi-view reconstruction; 3D-VAE-GAN [<xref ref-type="bibr" rid="ref-28">28</xref>], a hybrid VAE&#x2013;GAN framework for volumetric generation; Differentiable Ray Consistency (DRC) [<xref ref-type="bibr" rid="ref-65">65</xref>], which leverages multi-view supervision; Pix2Mesh [<xref ref-type="bibr" rid="ref-76">76</xref>], a mesh-based generative model; and Occupancy Networks (ONet) [<xref ref-type="bibr" rid="ref-49">49</xref>], which represent 3D shapes as continuous implicit functions.</p>
<p>In addition, we include more recent voxel-based approaches such as TMVNet [<xref ref-type="bibr" rid="ref-40">40</xref>], which utilizes a Transformer-based 3D encoder for modeling long-range volumetric dependencies, and Pix2Vox-A [<xref ref-type="bibr" rid="ref-39">39</xref>], which adopts a context-aware multi-view voxel refinement strategy.</p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> provides a detailed comparison of SWAGE-3D with these methods based on category-specific and average Intersection over Union (IoU) scores evaluated on the ShapeNet dataset. Although TMVNet and Pix2Vox-A achieve higher overall IoU values on ShapeNet, it is important to contextualize these results within architectural and training paradigm differences. Both TMVNet and Pix2Vox-A are designed around multi-view supervision and feature fusion mechanisms, enabling stronger geometric consistency during training. In contrast, SWAGE-3D operates under a single-view generative framework enhanced with depth integration and adversarial learning.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of different 3D model generation methods based on IoU metric.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th align="center" colspan="8">Method</th>
</tr>
<tr>
<th></th>
<th>3D-R2N2</th>
<th>3DVAEGAN</th>
<th>DRC</th>
<th>Pix2Mesh</th>
<th>ONet</th>
<th>TMVNet</th>
<th>Pix2Vox-A</th>
<th>Ours</th>
</tr>
<tr>
<th></th>
<th>[<xref ref-type="bibr" rid="ref-30">30</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-28">28</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-65">65</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-76">76</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-49">49</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-40">40</xref>]</th>
<th>[<xref ref-type="bibr" rid="ref-39">39</xref>]</th>
<th>&#x2013;</th>
</tr>
</thead>
<tbody>
<tr>
<td>Airplane</td>
<td>0.513</td>
<td>0.420</td>
<td>0.571</td>
<td>0.420</td>
<td>0.571</td>
<td>0.691</td>
<td>0.684</td>
<td>0.572</td>
</tr>
<tr>
<td>Bench</td>
<td>0.421</td>
<td>0.340</td>
<td>0.453</td>
<td>0.323</td>
<td>0.485</td>
<td>0.659</td>
<td>0.616</td>
<td>0.401</td>
</tr>
<tr>
<td>Cabinet</td>
<td>0.716</td>
<td>0.600</td>
<td>0.635</td>
<td>0.664</td>
<td>0.733</td>
<td>0.853</td>
<td>0.792</td>
<td>0.726</td>
</tr>
<tr>
<td>Car</td>
<td>0.798</td>
<td>0.760</td>
<td>0.755</td>
<td>0.552</td>
<td>0.737</td>
<td>0.870</td>
<td>0.854</td>
<td>0.835</td>
</tr>
<tr>
<td>Chair</td>
<td>0.466</td>
<td>0.360</td>
<td>0.469</td>
<td>0.396</td>
<td>0.501</td>
<td>0.721</td>
<td>0.567</td>
<td>0.460</td>
</tr>
<tr>
<td>Display</td>
<td>0.468</td>
<td>0.400</td>
<td>0.419</td>
<td>0.490</td>
<td>0.471</td>
<td>0.595</td>
<td>0.537</td>
<td>0.419</td>
</tr>
<tr>
<td>Lamp</td>
<td>0.381</td>
<td>0.320</td>
<td>0.415</td>
<td>0.323</td>
<td>0.371</td>
<td>0.534</td>
<td>0.443</td>
<td>0.423</td>
</tr>
<tr>
<td>Loudspeaker</td>
<td>0.662</td>
<td>0.590</td>
<td>0.609</td>
<td>0.599</td>
<td>0.647</td>
<td>0.712</td>
<td>0.714</td>
<td>0.693</td>
</tr>
<tr>
<td>Rifle</td>
<td>0.544</td>
<td>0.540</td>
<td>0.608</td>
<td>0.402</td>
<td>0.474</td>
<td>0.783</td>
<td>0.615</td>
<td>0.463</td>
</tr>
<tr>
<td>Sofa</td>
<td>0.628</td>
<td>0.570</td>
<td>0.606</td>
<td>0.613</td>
<td>0.680</td>
<td>0.701</td>
<td>0.709</td>
<td>0.662</td>
</tr>
<tr>
<td>Table</td>
<td>0.513</td>
<td>0.330</td>
<td>0.424</td>
<td>0.395</td>
<td>0.506</td>
<td>0.660</td>
<td>0.601</td>
<td>0.524</td>
</tr>
<tr>
<td>Telephone</td>
<td>0.661</td>
<td>0.680</td>
<td>0.413</td>
<td>0.661</td>
<td>0.720</td>
<td>0.801</td>
<td>0.776</td>
<td>0.731</td>
</tr>
<tr>
<td>Watercraft</td>
<td>0.513</td>
<td>0.480</td>
<td>0.556</td>
<td>0.397</td>
<td>0.530</td>
<td>0.685</td>
<td>0.594</td>
<td>0.544</td>
</tr>
<tr>
<td><bold>Average</bold></td>
<td>0.560</td>
<td>0.491</td>
<td>0.533</td>
<td>0.479</td>
<td>0.571</td>
<td>0.712</td>
<td>0.661</td>
<td>0.573</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>TMVNet leverages a Transformer-based 3D encoder to model long-range volumetric dependencies across multiple views, which provides stronger structural priors during reconstruction. Similarly, Pix2Vox-A employs a context-aware refinement strategy that explicitly aggregates multi-view voxel predictions, leading to improved geometric completeness.</p>
<p>By comparison, SWAGE-3D prioritizes generative modeling stability and distribution learning through a hybrid VAE-WGAN framework. While adversarial learning enhances structural realism and diversity, it does not directly optimize IoU in a purely supervised regression sense. Therefore, the slightly lower IoU values should be interpreted as a trade-off between generative robustness and direct voxel regression performance rather than a deficiency in reconstruction capability.</p>
<p>The results reveal a competitive landscape where different methods excel in specific categories. Across categories, SWAGE-3D consistently improves over the baseline 3D-VAE-GAN and achieves its most pronounced gains on rigid object classes (e.g., Car and Telephone), suggesting that the proposed stabilization and depth fusion improve reconstruction reliability. For instance, among single-view baselines, SWAGE-3D demonstrates a substantial margin in the Car category, achieving an IoU score of 0.835, and remains highly competitive even against multi-view methods. This exceptional performance in rigid, man-made categories is largely driven by the depth-aware encoder, which successfully resolves the depth ambiguity of flat surfaces that plagues purely RGB-based moels. While multi-view architectures like TMVNet establish the upper bound in categories such as Bench, Cabinet, and Rifle, SWAGE-3D achieves the highest average IoU among the compared single-view generative voxel-based baselines. This underscores the framework&#x2019;s ability to deliver high-quality reconstructions across a wide range of object geometries, offering a robust balance between detail and structural coherence.</p>
<p>It is worth noting that while implicit representation methods like ONet excel in organic shapes due to their continuous nature, SWAGE-3D demonstrates superior performance in rigid, geometric objects (e.g., Cars with 0.835 IoU vs. ONet&#x2019;s 0.737). This indicates that our voxel-based approach, fortified with depth guidance and attention, is particularly effective for preserving the structural integrity of man-made objects with defined planar surfaces.</p>
<p>While the proposed SWAGE-3D model demonstrates superior performance in volumetric objects such as cars and loudspeakers, a slight performance drop is observed in categories characterized by thin and complex structures, such as Rifle and Bench. This limitation is attributed to the fixed voxel resolution (<inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msup><mml:mn>32</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula>) and the regularization effects of spectral normalization, which may occasionally treat fine structural details as high-frequency noise during the reconstruction process. Increasing the voxel resolution in future works could effectively mitigate this limitation.</p>
<p>To further validate the effectiveness of SWAGE-3D, we compare its performance with the baseline 3DVAEGAN model using the IoU metric, as illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. When evaluating individual objects, SWAGE-3D demonstrates significant improvements. Specifically, it achieves a 15.2% higher IoU for the Airplane category and a 19.4% improvement for the Table category compared to 3DVAEGAN. For objects like Display and Telephone, SWAGE-3D shows modest gains of 1.9% and 5.1%, respectively. However, in the Rifle category, 3DVAEGAN outperforms SWAGE-3D by 7.7%. Compared to the baseline 3D-VAE-GAN evaluated under the same single-view voxel reconstruction protocol, SWAGE-3D improves the average IoU from 0.491 to 0.573, corresponding to a 16.7% relative gain.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparison of 3DVAEGAN and SWAGE-3D models on IoU metric for 13 objects.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-6.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> provides qualitative insights into the performance of SWAGE-3D. The figure illustrates the input RGB image, the associated depth map (colorized for enhanced visibility), the voxelized model generated by SWAGE-3D, the corresponding ground truth voxelized model from the dataset, and the intersection volume between the predicted and ground truth voxels. To highlight performance variations, examples representing the highest (Car), average (Airplane), and lowest (Bench) IoU values achieved by SWAGE-3D are included. These qualitative results corroborate the quantitative findings, demonstrating the framework&#x2019;s ability to produce highly accurate and structurally coherent 3D models.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Qualitative performance of SWAGE-3D: input RGB and depth maps (colorized), predicted voxels, ground truth voxels, and their intersection for high (Car), average (Airplane), and low (Bench) IoU results.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-7.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-8">Fig. 8</xref> provides a qualitative and quantitative comparison of 3D-VAE-GAN, 3D-R2N2, Pix2Mesh, and SWAGE-3D. Depth maps are included in the inputs section but are only utilized for SWAGE-3D. Ground Truth refers to the actual 3D voxelized models in the dataset. Analysis of the 3D results highlights the superior performance of SWAGE-3D. Similar to the qualitative results presented earlier, SWAGE-3D demonstrates strong performance for Airplane, Car, Lamp, Loudspeaker, Table, and Telephone. For the Airplane, it produces a model very similar to the real one, though minor extraneous voxels appear outside the structure. For the Bench, while there are deficiencies in the railing parts, the basic skeleton is correctly constructed.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Comparison of 3D-VAE-GAN, 3D-R2N2, Pix2Mesh, and SWAGE-3D across object categories.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_79254-fig-8.tif"/>
</fig>
<p>For the Cabinet, despite some issues with the arm sections, the fundamental structure is accurately produced. For the Car, the generated model closely resembles the real object. For the Chair, errors exist in the backrest and seating areas, but the overall production is successful. For the Display, the model includes the stand and is well-produced. The Lamp is reconstructed with high success. For the Loudspeaker, despite minor errors in the front section, the output is satisfactory. For the Rifle, intermediate voxels are missing. This limitation is attributed to the fixed voxel resolution (<inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msup><mml:mn>32</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula>) and the aggressive regularization of Spectral Normalization, which occasionally treats very thin structures as high-frequency noise. However, the model still correctly localizes the main body, suggesting that higher resolutions in future work could resolve this. The Sofa is produced with great success, with only minor voxel deficiencies in the seating area. For the Table, the generated model closely matches the real one, though minor errors persist in the foot sections. The Telephone is reconstructed with 100% accuracy. For the Watercraft, the model performs well, despite some deficiencies in the front and upper sections. Overall, SWAGE-3D achieves excellent results compared to other models in the literature, particularly excelling in specific object categories.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Incremental Ablation Study</title>
<p>To systematically quantify the individual contributions of the integrated techniques within the SWAGE-3D framework, an incremental ablation study was conducted. Due to the high computational demand of full-scale training, this study was performed on six representative categories (Airplane, Bench, Car, Lamp, Loudspeaker, and Telephone) that exhibit diverse geometric characteristics. To ensure a fair and rigorous evaluation, all ablation variants were trained under a separate, strictly unified experimental setup to isolate the relative gain of each component. While absolute values may vary slightly from the peak performance results reported in <xref ref-type="table" rid="table-3">Table 3</xref> due to the stochastic nature of adversarial training, the performance trends across configurations remain entirely consistent.</p>

<p>Five model variants were evaluated: (1) a baseline 3D-VAE-GAN using 3-channel RGB input, (2) the integration of monocular depth maps (RGB-D), (3) the addition of Wasserstein GAN loss with Gradient Penalty (WGAN-GP), (4) the inclusion of Self-Attention modules, and (5) the final weighted ensemble inference. The results across Mean IoU, F-Score, and Chamfer Distance (CD) are summarized in <xref ref-type="table" rid="table-4">Tables 4</xref>&#x2013;<xref ref-type="table" rid="table-6">6</xref>, respectively.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Incremental ablation study results for Mean IoU across six representative categories.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center">Configuration</th>
<th align="center">Airplane</th>
<th align="center">Bench</th>
<th align="center">Car</th>
<th align="center">Lamp</th>
<th align="center">Loudspeaker</th>
<th align="center">Telephone</th>
<th align="center">Mean</th>
</tr>
</thead>
<tbody>
<tr>
<td>(1) Baseline (RGB)</td>
<td>0.557</td>
<td>0.380</td>
<td>0.828</td>
<td>0.340</td>
<td>0.677</td>
<td>0.729</td>
<td>0.585</td>
</tr>
<tr>
<td>(2) &#x002B; Depth Map</td>
<td>0.570</td>
<td>0.385</td>
<td>0.825</td>
<td><bold>0.431</bold></td>
<td>0.689</td>
<td><bold>0.730</bold></td>
<td><bold>0.605</bold></td>
</tr>
<tr>
<td>(3) &#x002B; WGAN-GP</td>
<td><bold>0.581</bold></td>
<td>0.387</td>
<td><bold>0.829</bold></td>
<td>0.395</td>
<td>0.689</td>
<td>0.727</td>
<td>0.601</td>
</tr>
<tr>
<td>(4) &#x002B; Attention</td>
<td>0.568</td>
<td><bold>0.392</bold></td>
<td>0.812</td>
<td>0.410</td>
<td><bold>0.691</bold></td>
<td>0.717</td>
<td>0.598</td>
</tr>
<tr>
<td>(5) &#x002B; Ensemble</td>
<td>0.567</td>
<td><bold>0.392</bold></td>
<td>0.816</td>
<td>0.412</td>
<td>0.688</td>
<td>0.715</td>
<td>0.598</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-4fn1" fn-type="other">
<p>Note: Bold values indicate the best performance.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Incremental ablation study results for F-Score across six representative categories.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center">Configuration</th>
<th align="center">Airplane</th>
<th align="center">Bench</th>
<th align="center">Car</th>
<th align="center">Lamp</th>
<th align="center">Loudspeaker</th>
<th align="center">Telephone</th>
<th align="center">Mean</th>
</tr>
</thead>
<tbody>
<tr>
<td>(1) Baseline (RGB)</td>
<td>0.715</td>
<td>0.550</td>
<td><bold>0.906</bold></td>
<td>0.506</td>
<td>0.807</td>
<td><bold>0.843</bold></td>
<td>0.721</td>
</tr>
<tr>
<td>(2) &#x002B; Depth Map</td>
<td>0.725</td>
<td>0.554</td>
<td>0.904</td>
<td><bold>0.598</bold></td>
<td>0.815</td>
<td><bold>0.843</bold></td>
<td><bold>0.740</bold></td>
</tr>
<tr>
<td>(3) &#x002B; WGAN-GP</td>
<td><bold>0.734</bold></td>
<td>0.555</td>
<td><bold>0.906</bold></td>
<td>0.563</td>
<td>0.816</td>
<td>0.841</td>
<td>0.736</td>
</tr>
<tr>
<td>(4) &#x002B; Attention</td>
<td>0.724</td>
<td><bold>0.562</bold></td>
<td>0.896</td>
<td>0.580</td>
<td><bold>0.817</bold></td>
<td>0.834</td>
<td>0.735</td>
</tr>
<tr>
<td>(5) &#x002B; Ensemble</td>
<td>0.723</td>
<td>0.561</td>
<td>0.899</td>
<td>0.579</td>
<td>0.815</td>
<td>0.833</td>
<td>0.735</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p>Note: Bold values indicate the best performance.</p>
</fn>
</table-wrap-foot>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Incremental ablation study results for Chamfer Distance (CD).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center">Configuration</th>
<th align="center">Airplane</th>
<th align="center">Bench</th>
<th align="center">Car</th>
<th align="center">Lamp</th>
<th align="center">Speaker</th>
<th align="center">Telephone</th>
<th align="center">Mean</th>
</tr>
</thead>
<tbody>
<tr>
<td>(1) Baseline (RGB)</td>
<td>0.955</td>
<td>1.618</td>
<td><bold>0.214</bold></td>
<td>10.015</td>
<td>0.925</td>
<td>0.423</td>
<td>2.358</td>
</tr>
<tr>
<td>(2) &#x002B; Depth Map</td>
<td>0.885</td>
<td>1.597</td>
<td>0.219</td>
<td>4.157</td>
<td>0.874</td>
<td><bold>0.407</bold></td>
<td>1.357</td>
</tr>
<tr>
<td>(3) &#x002B; WGAN-GP</td>
<td>0.874</td>
<td>1.825</td>
<td>0.219</td>
<td>5.841</td>
<td>0.841</td>
<td><bold>0.399</bold></td>
<td>1.667</td>
</tr>
<tr>
<td>(4) &#x002B; Attention</td>
<td>0.845</td>
<td><bold>1.469</bold></td>
<td>0.233</td>
<td>4.252</td>
<td><bold>0.839</bold></td>
<td>0.413</td>
<td>1.342</td>
</tr>
<tr>
<td>(5) &#x002B; Ensemble</td>
<td><bold>0.826</bold></td>
<td>1.499</td>
<td>0.227</td>
<td><bold>4.051</bold></td>
<td>0.872</td>
<td>0.423</td>
<td><bold>1.316</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-6fn1" fn-type="other">
<p>Note: Bold values indicate the best performance.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The systematic ablation reveals that the integration of depth maps (Config 2) is the primary driver for recovering geometric fidelity, resolving ambiguities where RGB data alone is insufficient. This is most evident in the Lamp category, where depth maps facilitated a substantial reduction in total geometric error. Subsequent architectural refinements via WGAN-GP and Self-Attention (Configs 3 and 4) focused on enhancing global shape consistency and boundary sharpness. While point-wise metrics like Mean IoU exhibit minor fluctuations, a common characteristic of GAN-based models where distributional realism is prioritized over local MSE, the structural metrics (F-Score and CD) indicate a consistent downward trend in geometric error across configurations. This study confirms that the cumulative success of SWAGE-3D arises from the synergistic combination of depth-assisted geometric initialization, stabilized adversarial learning, and robust ensemble inference.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Resolution Analysis (32<inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msup><mml:mi>32</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and 64<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:msup><mml:mi>64</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>)</title>
<p>To address the scalability of SWAGE-3D and evaluate its performance at higher voxel densities, we conducted a comparative analysis between the standard <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msup><mml:mn>32</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula> resolution and a higher <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msup><mml:mn>64</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula> grid size. This experiment was specifically performed on the Bench category, as it contains thin and complex structures that provide a rigorous test for high-resolution 3D generation. To ensure a fair assessment of computational efficiency and scaling behavior, both models were trained for 50 epochs under consistent experimental conditions.</p>
<p>The results, summarized in <xref ref-type="table" rid="table-7">Table 7</xref>, demonstrate that SWAGE-3D effectively scales to higher resolutions while maintaining structural integrity. While the <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msup><mml:mn>64</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula> resolution significantly increases the learning complexity and the state space of the voxel grid, the model successfully captures the essential geometric properties of the target shapes. As expected, the computational demand increases with resolution; specifically, GPU memory usage rose from 0.41 to 1.34 GB, and the training time per epoch increased from 1.08 to 11.25 min. Although the IoU and F-score values for the <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msup><mml:mn>64</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula> resolution are slightly lower than those of <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msup><mml:mn>32</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula> within the fixed 50-epoch training budget the model&#x2019;s ability to handle eight times the voxel density confirms its architectural robustness and scalability for high-fidelity 3D reconstruction.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Resolution analysis of SWAGE-3D under different voxel grid sizes. Performance is evaluated on the ShapeNet test split (Bench category) after 50 training epochs. Computational cost is measured on a single NVIDIA GTX 980 Ti GPU.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Resolution</th>
<th>IoU</th>
<th>F-Score (<inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula>)</th>
<th>Chamfer Dist. (<inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula>)</th>
<th>GPU Memory (GB)</th>
<th>Training Time (min)</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msup><mml:mn>32</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula></td>
<td>0.382</td>
<td>0.552</td>
<td>1.5416</td>
<td>0.41</td>
<td>1.08</td>
</tr>
<tr>
<td><inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msup><mml:mn>64</mml:mn><mml:mn>3</mml:mn></mml:msup></mml:math></inline-formula></td>
<td>0.343</td>
<td>0.506</td>
<td>3.0343</td>
<td>1.34</td>
<td>11.25</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Discussion</title>
<p>The experimental findings provide a comprehensive assessment of the proposed SWAGE-3D framework across reconstruction accuracy, geometric consistency, stabilization behavior, and scalability. The results collectively indicate that the integration of depth-assisted encoding and stabilized adversarial training contributes to improved reconstruction robustness under a single-view voxel-based setting.</p>
<p>The quantitative comparison on ShapeNet demonstrates that SWAGE-3D consistently outperforms the baseline 3D-VAE-GAN across multiple object categories in terms of IoU. The most pronounced improvements are observed in rigid and structurally well-defined categories such as Car and Telephone, suggesting that depth-guided encoding effectively mitigates depth ambiguity in planar and symmetric structures. This indicates that the RGB-D fusion strategy plays a central role in improving geometric reliability when only a single image is available.</p>
<p>The incremental ablation study further clarifies the relative contributions of architectural components. The depth map integration yields the most substantial improvement in mean IoU and F-score, confirming that geometric cues extracted from monocular depth estimation significantly enhance volumetric inference. The inclusion of WGAN-GP and spectral normalization primarily contributes to training stabilization and distribution alignment rather than dramatic IoU gains. While the improvement in volumetric overlap metrics is moderate, the reduction in Chamfer Distance suggests improved geometric smoothness and surface coherence. Similarly, the self-attention module provides marginal gains in voxel overlap but contributes to enhanced structural consistency, particularly reflected in distance-based evaluation. The ensemble strategy offers limited incremental improvement in IoU but helps reduce variance and improve robustness during inference.</p>
<p>Taken together, these findings suggest that performance gains do not stem from a single dominant architectural modification alone, but rather from a controlled combination of complementary stabilization and geometric enhancement strategies. The improvements are more pronounced in structural reliability and geometric consistency than in raw volumetric overlap alone.</p>
<p>The inclusion of F-score and Chamfer Distance provides a more nuanced understanding of reconstruction behavior. While IoU captures volumetric overlap, F-score reflects precision-recall balance of occupancy prediction, and Chamfer Distance quantifies geometric proximity. The observed trend, where some configurations yield modest IoU changes but measurable Chamfer improvements, indicates that adversarial stabilization and attention mechanisms primarily refine surface quality rather than drastically altering voxel occupancy ratios. This highlights the importance of multi-metric evaluation when analyzing generative volumetric models.</p>
<p>The resolution analysis comparing 32<inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msup><mml:mi>32</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and 64<inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msup><mml:mi>64</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> voxel grids reveals a critical trade-off between geometric granularity and computational feasibility. Although the 64<inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msup><mml:mi>64</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> configuration increases representational capacity, it leads to reduced IoU and F-score under the current hardware constraints, while significantly increasing memory consumption and training time. The cubic growth in voxel space dramatically enlarges the optimization landscape, making stable adversarial training more challenging on a single GTX 980 Ti GPU. These findings indicate that while the architecture is technically scalable to higher resolutions, effective optimization at 64<inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msup><mml:mi>64</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> resolution requires either stronger regularization strategies or more advanced computational resources. Therefore, 32<inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msup><mml:mi>32</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> remains a practical and balanced resolution choice for stable single-view adversarial voxel reconstruction under moderate hardware settings.</p>
<p>Despite its improvements, SWAGE-3D has several limitations. First, the voxel-based representation inherently suffers from discretization artifacts and cubic computational complexity. Although adversarial learning improves structural realism, thin structures and fine details remain difficult to reconstruct at low resolutions. While higher resolutions are theoretically feasible, they introduce substantial computational overhead and optimization instability. Second, the model relies on monocular depth estimation as a preprocessing step. Errors in predicted depth maps directly propagate to the latent representation, potentially limiting reconstruction fidelity in visually ambiguous or textureless regions.</p>
<p>Third, while the incremental ablation study demonstrates the contribution of each component, the magnitude of improvement beyond depth integration is moderate. This suggests that geometric cues play a more dominant role than adversarial refinement in single-view voxel reconstruction under the current configuration. Finally, SWAGE-3D is evaluated under a single-view generative paradigm. Multi-view transformer-based methods achieve higher absolute IoU values by leveraging stronger geometric supervision. The proposed framework prioritizes stabilization and distribution-aware volumetric generation rather than purely supervised regression optimization.</p>
<p>Future research may explore hybrid voxel-implicit representations to alleviate discretization constraints, adaptive resolution strategies to dynamically allocate voxel density, and more advanced regularization techniques to stabilize higher-resolution adversarial training. Additionally, integrating uncertainty-aware depth estimation or jointly optimizing depth and voxel generation in an end-to-end manner may further enhance geometric consistency.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This study was motivated by persistent challenges in 3D object generation, particularly the limitations of existing 3D-VAE-GAN architectures in capturing fine-grained detail, ensuring training stability, and using spatial information from single-view inputs. To overcome these obstacles, we proposed SWAGE-3D, a novel framework that integrates several state-of-the-art enhancements including depth map integration, MixUp data augmentation, a ResNet18-based encoder, spectral normalization, self-attention mechanisms, Wasserstein GAN with gradient penalty (WGAN-GP), and ensemble model testing. These contributions collectively aim to address gaps in the literature by combining complementary techniques that enhance both the stability and fidelity of voxel-based 3D reconstructions. Our results confirm that while individual components like Attention or Depth Integration are powerful on their own, their combined application creates a necessary stability for voxel-based GANs, yielding a 16.7% improvement without requiring complex multi-stage training pipelines.</p>
<p>In this study, we introduced SWAGE-3D, an advanced framework designed to overcome the limitations of existing 3D generative architectures by integrating depth-enhanced input, self-attention mechanisms, ensemble learning, and training stabilization techniques such as spectral normalization and WGAN-GP. Unlike prior studies that focused primarily on the standard 3D-VAE-GAN, our work systematically incorporates and improves upon a broader range of recent state-of-the-art (SOTA) models, offering a more robust and accurate solution for single-view 3D object reconstruction.</p>
<p>Extensive evaluations were conducted on the ShapeNet dataset using the Intersection over Union (IoU) metric. Compared to the baseline 3D-VAE-GAN under identical single-view voxel reconstruction settings, SWAGE-3D achieves a relative IoU improvement of 16.7%, increasing the average IoU from 0.491 to 0.573. Beyond this, while multi-view architectures achieve higher overall bounds, SWAGE-3D demonstrated superior or comparable performance when benchmarked against leading single-view baselines such as ONet (0.571), 3D-R2N2 (0.560), and DRC (0.533). These results highlight SWAGE-3D&#x2019;s capacity to produce structurally coherent, detail-rich 3D reconstructions.</p>
<p>This performance gain can be attributed to several architectural and training innovations. The integration of Depth Anything V2-based depth maps as a fourth input channel enhanced the geometric reasoning of the encoder. The ResNet18-based encoder, adapted for 4-channel input, enabled efficient and accurate feature extraction. The attention-augmented generator improved global context understanding, while the WGAN-GP &#x002B; spectral normalization combination addressed training instability, leading to smoother convergence. Moreover, ensemble testing across selected model checkpoints further improved inference robustness and generalization.</p>
<p>Looking ahead, future research directions include extending SWAGE-3D to handle real-world noisy or occluded inputs, and exploring alternative representations such as point clouds or implicit surfaces to improve scalability and detail fidelity. Additionally, incorporating multi-view fusion, cross-modal learning, and transformer-based encoders may further elevate reconstruction accuracy and enable broader applicability across domains like autonomous navigation, AR/VR, and medical imaging.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Zafer Serin: Conceptualization, Methodology, Software, Validation, Visualization, Writing&#x2014;Original Draft; Cihan Karakuzu: Conceptualization, Supervision, Writing&#x2014;Review &#x0026; Editing; U&#x011F;ur Y&#x00FC;zge&#x00E7;: Conceptualization, Methodology, Supervision, Writing&#x2014;Review &#x0026; Editing. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The experiments in this study were conducted using the ShapeNetCore dataset, publicly available at the official Stanford repository (<ext-link ext-link-type="uri" xlink:href="https://shapenet.org/">https://shapenet.org/</ext-link>). Access to the dataset requires registration through the official website. In this work, we utilized a subset of 13 object categories following the common experimental protocol adopted in prior voxel-based reconstruction studies. The source code, pre-trained models, and detailed instructions to reproduce the results will be made publicly available on GitHub upon acceptance of the manuscript.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name></person-group>. <article-title>DGGR-Net: single-image 3D reconstruction from complex backgrounds via graph-based refinement and difference-guided fusion</article-title>. <source>J King Saud Univ Comput Inf Sci</source>. <year>2025</year>;<volume>37</volume>:<fpage>222</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s44443-025-00251-8</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kazhdan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hoppe</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Screened poisson surface reconstruction</article-title>. <source>ACM Trans Graph</source>. <year>2013</year>;<volume>32</volume>(<issue>3</issue>):<fpage>29</fpage>. doi:<pub-id pub-id-type="doi">10.1145/2487228.2487237</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rusu</surname> <given-names>RB</given-names></string-name>, <string-name><surname>Cousins</surname> <given-names>S</given-names></string-name></person-group>. <article-title>3D is here: point cloud library (PCL)</article-title>. In: <conf-name>2011 IEEE International Conference on Robotics and Automation (ICRA)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2011</year>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icra.2011.5980567</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Botsch</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kobbelt</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A remeshing approach to multiresolution modeling</article-title>. In: <conf-name>Proceedings of the 2004 Eurographics/ACM SIGGRAPH Symposium on Geometry Processing</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>ACM</publisher-name>; <year>2004</year>. p. <fpage>185</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1145/1057432.1057457</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>LeCun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Haffner</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc IEEE</source>. <year>1998</year>;<volume>86</volume>(<issue>11</issue>):<fpage>2278</fpage>&#x2013;<lpage>324</lpage>. doi:<pub-id pub-id-type="doi">10.1109/5.726791</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name></person-group>. <chapter-title>ImageNet classification with deep convolutional neural networks</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2012</year>. p. <fpage>1097</fpage>&#x2013;<lpage>105</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <comment>arXiv:1409.1556. 2015</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shrivastava</surname> <given-names>A</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Revisiting unreasonable effectiveness of data in deep learning era</article-title>. In: <conf-name>2017 IEEE International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>843</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2017.97</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Cisse</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dauphin</surname> <given-names>YN</given-names></string-name>, <string-name><surname>Lopez-Paz</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Mixup: beyond empirical risk minimization</article-title>. <comment>arXiv:1710.09412. 2018</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational Bayes</article-title>. <comment>arXiv:1312.6114. 2014</comment>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Goodfellow</surname> <given-names>I</given-names></string-name>, <string-name><surname>Pouget-Abadie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Mirza</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Warde-Farley</surname> <given-names>D</given-names></string-name>, <string-name><surname>Ozair</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>Generative adversarial nets</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2014</year>. p. <fpage>2672</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chakraborty</surname> <given-names>T</given-names></string-name>, <string-name><surname>Reddy</surname> <given-names>KSU</given-names></string-name>, <string-name><surname>Naik</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Panja</surname> <given-names>M</given-names></string-name>, <string-name><surname>Manvitha</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Ten years of generative adversarial nets (GANs): a survey of the state-of-the-art</article-title>. <source>Mach Learn Sci Technol</source>. <year>2024</year>;<volume>5</volume>(<issue>1</issue>):<fpage>011001</fpage>. doi:<pub-id pub-id-type="doi">10.1088/2632-2153/ad1f77</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Arjovsky</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chintala</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Wasserstein generative adversarial networks</article-title>. In: <conf-name>Proceedings of the 34th International Conference on Machine Learning</conf-name>. <publisher-loc>Brookline, MA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2017</year>. p. <fpage>214</fpage>&#x2013;<lpage>23</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>Attention is all you need</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017</year>. p. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bahdanau</surname> <given-names>D</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Neural machine translation by jointly learning to align and translate</article-title>. In: <conf-name>Proceeding of the International Conference on Learning Representations (ICLR)</conf-name>; <year>2015 May 7&#x2013;9</year>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>R</given-names></string-name>, <string-name><surname>Van Hoorick</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tokmakov</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zakharov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Vondrick</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Zero-1-to-3: zero-shot one image to 3D object</article-title>. In: <conf-name>2023 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>9264</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv51070.2023.00853</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Poole</surname> <given-names>B</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Barron</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Mildenhall</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Dreamfusion: text-to-3D using 2D diffusion</article-title>. <comment>arXiv:2209.14988. 2022</comment>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nichol</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Shap-E: generating conditional 3D implicit functions</article-title>. <comment>arXiv:2305.02463. 2023</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nichol</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Mishkin</surname> <given-names>P</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Point-E: a system for generating 3D point clouds from complex prompts</article-title>. <comment>arXiv:2212.08751. 2022</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Smith</surname> <given-names>LN</given-names></string-name></person-group>. <article-title>Cyclical learning rates for training neural networks</article-title>. In: <conf-name>2017 IEEE Winter Conference on Applications of Computer Vision (WACV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>464</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1109/wacv.2017.58</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Loshchilov</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hutter</surname> <given-names>F</given-names></string-name></person-group>. <article-title>SGDR: stochastic gradient descent with warm restarts</article-title>. In: <conf-name>ICLR 2017 (5th International Conference on Learning Representations)</conf-name>; <year>2017 Apr 24&#x2013;26</year>; <publisher-loc>Toulon, France</publisher-loc>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dietterich</surname> <given-names>TG</given-names></string-name></person-group>. <article-title>Ensemble methods in machine learning</article-title>. In: <conf-name>Proceedings of the First International Workshop on Multiple Classifier Systems</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2000</year>. p. <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1007/3-540-45014-9_1</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sagi</surname> <given-names>O</given-names></string-name>, <string-name><surname>Rokach</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Ensemble learning: a survey</article-title>. <source>Wiley Interdiscip Rev Data Min Knowl Discov</source>. <year>2018</year>;<volume>8</volume>(<issue>4</issue>):<fpage>e1249</fpage>. doi:<pub-id pub-id-type="doi">10.1002/widm.1249</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Striuk</surname> <given-names>O</given-names></string-name>, <string-name><surname>Kondratenko</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Optimization strategy for generative adversarial networks design</article-title>. <source>Int J Comput</source>. <year>2023</year>;<volume>22</volume>(<issue>3</issue>):<fpage>292</fpage>&#x2013;<lpage>301</lpage>. doi:<pub-id pub-id-type="doi">10.47839/ijc.22.3.3223</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>DY</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>XP</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>YT</given-names></string-name>, <string-name><surname>Ouhyoung</surname> <given-names>M</given-names></string-name></person-group>. <article-title>On visual similarity based 3D model retrieval</article-title>. <source>Comput Graph Forum</source>. <year>2003</year>;<volume>22</volume>(<issue>3</issue>):<fpage>223</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1111/1467-8659.00669</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Funkhouser</surname> <given-names>T</given-names></string-name>, <string-name><surname>Min</surname> <given-names>P</given-names></string-name>, <string-name><surname>Kazhdan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Halderman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dobkin</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A search engine for 3D models</article-title>. <source>ACM Trans Graph</source>. <year>2003</year>;<volume>22</volume>(<issue>1</issue>):<fpage>83</fpage>&#x2013;<lpage>105</lpage>. doi:<pub-id pub-id-type="doi">10.1145/588272.588279</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>T</given-names></string-name>, <string-name><surname>Freeman</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tenenbaum</surname> <given-names>J</given-names></string-name></person-group>. <chapter-title>Learning a probabilistic latent space of object shapes via 3D generative-adversarial modeling</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2016</year>. p. <fpage>82</fpage>&#x2013;<lpage>90</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>UniRecGen: unifying multi-view 3D reconstruction and generation</article-title>. <comment>arXiv:2604.01479. 2026</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Choy</surname> <given-names>CB</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>D</given-names></string-name>, <string-name><surname>Gwak</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Savarese</surname> <given-names>S</given-names></string-name></person-group>. <article-title>3D-R2N2: a unified approach for single and multi-view 3D object reconstruction</article-title>. In: <conf-name>The European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2016</year>. p. <fpage>628</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-46484-8_38</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Smith</surname> <given-names>EJ</given-names></string-name>, <string-name><surname>Meger</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Improved adversarial systems for 3D object generation and reconstruction</article-title>. In: <conf-name>Proceeding of the Conference on Robot Learning (CoRL)</conf-name>. <publisher-loc>Brookline, MA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2017</year>. p. <fpage>87</fpage>&#x2013;<lpage>96</lpage>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chan</surname> <given-names>ER</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>CZ</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Nagano</surname> <given-names>K</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>B</given-names></string-name>, <string-name><surname>De Mello</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Efficient geometry-aware 3D generative adversarial networks</article-title>. In: <conf-name>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>16102</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr52688.2022.01565</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Maturana</surname> <given-names>D</given-names></string-name>, <string-name><surname>Scherer</surname> <given-names>S</given-names></string-name></person-group>. <article-title>VoxNet: a 3D convolutional neural network for real-time object recognition</article-title>. In: <conf-name>2015 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>922</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iros.2015.7353481</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>H</given-names></string-name>, <string-name><surname>Maji</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kalogerakis</surname> <given-names>E</given-names></string-name>, <string-name><surname>Learned-Miller</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Multi-view convolutional neural networks for 3D shape recognition</article-title>. In: <conf-name>2015 IEEE International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2015</year>. p. <fpage>945</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2015.114</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Johns</surname> <given-names>E</given-names></string-name>, <string-name><surname>Leutenegger</surname> <given-names>S</given-names></string-name>, <string-name><surname>Davison</surname> <given-names>AJ</given-names></string-name></person-group>. <article-title>Pairwise decomposition of image sequences for active multi-view recognition</article-title>. In: <conf-name>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2016</year>. p. <fpage>3813</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2016.414</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hegde</surname> <given-names>V</given-names></string-name>, <string-name><surname>Zadeh</surname> <given-names>R</given-names></string-name></person-group>. <article-title>FusionNet: 3D object classification using multiple data representations</article-title>. <comment>arXiv:1607.05695. 2016</comment>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Brock</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lim</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ritchie</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Weston</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Generative and discriminative voxel modeling with convolutional neural networks</article-title>. <comment>arXiv:1608.04236. 2016</comment>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Sedaghat</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zolfaghari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Amiri</surname> <given-names>E</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Orientation-boosted voxel nets for 3D object recognition</article-title>. <comment>arXiv:1604.03351. 2016</comment>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Pix2Vox: context-aware 3D reconstruction from single and multi-view images</article-title>. In: <conf-name>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>2690</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv.2019.00278</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Peng</surname> <given-names>K</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Quarles</surname> <given-names>J</given-names></string-name>, <string-name><surname>Desai</surname> <given-names>K</given-names></string-name></person-group>. <article-title>TMVNet: using transformers for multi-view voxel-based 3D reconstruction</article-title>. In: <conf-name>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2022</year>. p. <fpage>221</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvprw56347.2022.00036</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Choy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Alvarez</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Fidler</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>VoxFormer: sparse voxel transformer for camera-based 3D semantic scene completion</article-title>. In: <conf-name>2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>9087</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00877</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shoukat</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Sargano</surname> <given-names>AB</given-names></string-name>, <string-name><surname>Malyshev</surname> <given-names>A</given-names></string-name>, <string-name><surname>You</surname> <given-names>L</given-names></string-name>, <string-name><surname>Habib</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>SS3DNet-AF: a single-stage, single-view 3D reconstruction network with attention-based fusion</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>23</issue>):<fpage>11424</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app142311424</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>K</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>I</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Park</surname> <given-names>U</given-names></string-name></person-group>. <article-title>Multi-head attention refiner for multi-view 3D reconstruction</article-title>. <source>J Imaging</source>. <year>2024</year>;<volume>10</volume>(<issue>11</issue>):<fpage>268</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jimaging10110268</pub-id>; <pub-id pub-id-type="pmid">39590732</pub-id></mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xiong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>3D voxel reconstruction from single-view image based on cross-domain feature fusion</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>256</volume>(<issue>1</issue>):<fpage>124957</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2024.124957</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Benes</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>SVDTree: semantic voxel diffusion for single image tree reconstruction</article-title>. In: <conf-name>2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2024</year>. p. <fpage>4692</fpage>&#x2013;<lpage>702</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr52733.2024.00449</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Su</surname> <given-names>H</given-names></string-name>, <string-name><surname>Guibas</surname> <given-names>LJ</given-names></string-name></person-group>. <article-title>A point set generation network for 3D object reconstruction from a single image</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>2463</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2017.264</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Achlioptas</surname> <given-names>P</given-names></string-name>, <string-name><surname>Diamanti</surname> <given-names>O</given-names></string-name>, <string-name><surname>Mitliagkas</surname> <given-names>I</given-names></string-name>, <string-name><surname>Guibas</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Learning representations and generative models for 3D point clouds</article-title>. In: <conf-name>The International Conference on Machine Learning (ICML)</conf-name>. <publisher-loc>Brookline, MA, USA</publisher-loc>: <publisher-name>PMLR</publisher-name>; <year>2018</year>. p. <fpage>40</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>YN</given-names></string-name></person-group>. <article-title>Learning descriptor networks for 3D shape synthesis and analysis</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>8629</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2018.00900</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mescheder</surname> <given-names>L</given-names></string-name>, <string-name><surname>Oechsle</surname> <given-names>M</given-names></string-name>, <string-name><surname>Niemeyer</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nowozin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Geiger</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Occupancy networks: learning 3D reconstruction in function space</article-title>. In: <conf-name>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2019</year>. p. <fpage>4455</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2019.00459</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chibane</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alldieck</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pons-Moll</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Implicit functions in feature space for 3D shape reconstruction and completion</article-title>. In: <conf-name>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2020</year>. p. <fpage>6968</fpage>&#x2013;<lpage>79</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.00700</pub-id>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>M&#x00FC;ller</surname> <given-names>T</given-names></string-name>, <string-name><surname>Evans</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schied</surname> <given-names>C</given-names></string-name>, <string-name><surname>Keller</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Instant neural graphics primitives with a multiresolution hash encoding</article-title>. <source>ACM Trans Graph</source>. <year>2022</year>;<volume>41</volume>(<issue>4</issue>):<fpage>102</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3528223.3530127</pub-id>.</mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kerbl</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kopanas</surname> <given-names>G</given-names></string-name>, <string-name><surname>Leimk&#x00FC;hler</surname> <given-names>T</given-names></string-name>, <string-name><surname>Drettakis</surname> <given-names>G</given-names></string-name></person-group>. <article-title>3D Gaussian splatting for real-time radiance field rendering</article-title>. <source>ACM Trans Graph</source>. <year>2023</year>;<volume>42</volume>(<issue>4</issue>):<fpage>139</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3592433</pub-id>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ben-Hamu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Maron</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kezurer</surname> <given-names>I</given-names></string-name>, <string-name><surname>Avineri</surname> <given-names>G</given-names></string-name>, <string-name><surname>Lipman</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Multi-chart generative surface modeling</article-title>. <source>ACM Trans Graph</source>. <year>2018</year>;<volume>37</volume>(<issue>6</issue>):<fpage>215</fpage>. doi:<pub-id pub-id-type="doi">10.1145/3272127.3275052</pub-id>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Groueix</surname> <given-names>T</given-names></string-name>, <string-name><surname>Fisher</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>VG</given-names></string-name>, <string-name><surname>Russell</surname> <given-names>BC</given-names></string-name>, <string-name><surname>Aubry</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A papier-m&#x00E2;ch&#x00E9; approach to learning 3D surface generation</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>216</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2018.00030</pub-id>.</mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Litany</surname> <given-names>O</given-names></string-name>, <string-name><surname>Bronstein</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bronstein</surname> <given-names>M</given-names></string-name>, <string-name><surname>Makadia</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Deformable shape completion with graph convolutional autoencoders</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2018</year>. p. <fpage>1886</fpage>&#x2013;<lpage>95</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2018.00202</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bogo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Romero</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pons-Moll</surname> <given-names>G</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name></person-group>. <article-title>Dynamic FAUST: registering human bodies in motion</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>5573</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2017.591</pub-id>.</mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Du</surname> <given-names>S</given-names></string-name>, <string-name><surname>Davis</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Semantic parametric reshaping of human body models</article-title>. In: <conf-name>2014 2nd International Conference on 3D Vision</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2014</year>. p. <fpage>41</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/3dv.2014.47</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Girdhar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Fouhey</surname> <given-names>DF</given-names></string-name>, <string-name><surname>Rodriguez</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Learning a predictable and generative vector representation for objects</article-title>. In: <conf-name>The European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2016</year>. p. <fpage>484</fpage>&#x2013;<lpage>99</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-46466-4_29</pub-id>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sharma</surname> <given-names>A</given-names></string-name>, <string-name><surname>Grau</surname> <given-names>O</given-names></string-name>, <string-name><surname>Fritz</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Vconv-dae: deep volumetric shape learning without object labels</article-title>. In: <conf-name>The European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2016</year>. p. <fpage>236</fpage>&#x2013;<lpage>50</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-49409-8_20</pub-id>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>J</given-names></string-name>, <string-name><surname>Torr</surname> <given-names>PHS</given-names></string-name>, <string-name><surname>Koltun</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Point transformer</article-title>. In: <conf-name>2021 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>16239</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv48922.2021.01595</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LRM: large reconstruction model for single image to 3D</article-title>. <comment>arXiv:2311.04400. 2024</comment>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Kovachki</surname> <given-names>N</given-names></string-name>, <string-name><surname>Azizzadenesheli</surname> <given-names>K</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Bhattacharya</surname> <given-names>K</given-names></string-name>, <string-name><surname>Stuart</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fourier neural operator for parametric partial differential equations</article-title>. <comment>arXiv:2010.08895. 2021</comment>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kovachki</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Azizzadenesheli</surname> <given-names>K</given-names></string-name>, <string-name><surname>Bhattacharya</surname> <given-names>K</given-names></string-name>, <string-name><surname>Stuart</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Neural operator: learning maps between function spaces with applications to pdes</article-title>. <source>J Mach Learn Res</source>. <year>2023</year>;<volume>24</volume>(<issue>89</issue>):<fpage>1</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Komura</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Sparseneus: fast generalizable neural surface reconstruction from sparse views</article-title>. In: <conf-name>The European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>210</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-19824-3_13</pub-id>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tulsiani</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>T</given-names></string-name>, <string-name><surname>Efros</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Malik</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Multi-view supervision for single-view reconstruction via differentiable ray consistency</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>209</fpage>&#x2013;<lpage>17</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr.2017.30</pub-id>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Kang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>J</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>Depth anything V2</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2024</year>. p. <fpage>21875</fpage>&#x2013;<lpage>911</lpage>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kuang</surname> <given-names>P</given-names></string-name></person-group>. <article-title>3D-VRVT: 3D voxel reconstruction from a single image with vision transformer</article-title>. In: <conf-name>2021 International Conference on Culture-Oriented Science &#x0026; Technology (ICCST)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>343</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccst53801.2021.00078</pub-id>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>3D reconstruction for multi-view objects</article-title>. <source>Comput Electr Eng</source>. <year>2023</year>;<volume>106</volume>:<fpage>108567</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compeleceng.2022.108567</pub-id>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>YL</given-names></string-name>, <string-name><surname>Shuai</surname> <given-names>HH</given-names></string-name>, <string-name><surname>Tam</surname> <given-names>ZR</given-names></string-name>, <string-name><surname>Chiu</surname> <given-names>HY</given-names></string-name></person-group>. <article-title>Gradient normalization for generative adversarial networks</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2021</year>. p. <fpage>6353</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1109/iccv48922.2021.00631</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Serin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Y&#x00FC;zge&#x00E7;</surname> <given-names>U</given-names></string-name>, <string-name><surname>Karakuzu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Pre-trained variational autoencoder approaches for generating 3D objects from 2D images</article-title>. In: <conf-name>2nd International Congress of Electrical and Computer Engineering</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2023</year>. p. <fpage>87</fpage>&#x2013;<lpage>101</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-52760-9_7</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Gulrajani</surname> <given-names>I</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>F</given-names></string-name>, <string-name><surname>Arjovsky</surname> <given-names>M</given-names></string-name>, <string-name><surname>Dumoulin</surname> <given-names>V</given-names></string-name>, <string-name><surname>Courville</surname> <given-names>AC</given-names></string-name></person-group>. <chapter-title>Improved training of Wasserstein GANs</chapter-title>. In: <source>Advances in Neural Information Processing Systems (NeurIPS)</source>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>; <year>2017</year>. p. <fpage>5767</fpage>&#x2013;<lpage>77</lpage>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Towards generalized implementation of Wasserstein distance in GANs</article-title>. In: <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence (AAAI)</conf-name>. <publisher-loc>Palo Alto, CA, USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>; <year>2021</year>. p. <fpage>10514</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v35i12.17258</pub-id>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>An introduction to variational autoencoders</article-title>. <source>Found Trends Mach Learn</source>. <year>2019</year>;<volume>12</volume>(<issue>4</issue>):<fpage>307</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1561/2200000056</pub-id>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Lipschitz constrained GANs via boundedness and continuity</article-title>. <source>Neural Comput Appl</source>. <year>2020</year>;<volume>32</volume>(<issue>24</issue>):<fpage>18271</fpage>&#x2013;<lpage>83</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-020-04954-z</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>AX</given-names></string-name>, <string-name><surname>Funkhouser</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guibas</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hanrahan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>ShapeNet: an information-rich 3D model repository</article-title>. <comment>arXiv:1512.03012. 2015</comment>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>YG</given-names></string-name></person-group>. <article-title>Pixel2Mesh: generating 3D mesh models from single RGB images</article-title>. In: <conf-name>The European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2018</year>. p. <fpage>55</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-01252-6_4</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>