Score
Designs, implements, and evaluates latent diffusion models that perform probabilistic generation or reconstruction by running a diffusion process in a learned low-dimensional latent space produced by an autoencoder/variational encoder—this includes building the encoder/decoder pipeline, training the diffusion prior in latent space, and integrating conditioning and sampling strategies. Specialised variants cover 3D and structured latent diffusion, where the latent representation encodes volumetric or other spatial structure (e.g., tri-planar or patch-structured latents) so generated outputs are coherent in three dimensions or across the chosen structure.
This survey addresses key challenges in 3D vision—occlusion robustness, point cloud sparsity, density imbalance, and high-dimensional computational bottlenecks—across four core tasks: 3D generation, point cloud reconstruction, shape completion, and scene synthesis. Methodologically, it introduces the first unified taxonomy capturing paradigm evolution, integrating denoising diffusion probabilistic models (DDPMs), 3D conditional encoders, multi-view feature alignment, implicit neural representations (INRs), and multimodal (text/image) guidance. The work rigorously delineates current performance limits and standardizes evaluation benchmarks. Crucially, it identifies three viable technical pathways forward: efficient sampling strategies, lightweight backward processes, and large-scale 3D pretraining. These contributions provide both theoretical foundations and practical guidelines for advancing diffusion-based 3D modeling.
Diffusion models face challenges in image inverse problems, including difficulty in conditional sampling and latent-space encoder-decoder mismatch. To address these, this paper proposes a latent-space conditional sampling framework based on Sequential Monte Carlo (SMC). It is the first work to systematically integrate SMC into the latent space of diffusion models: auxiliary observations are injected during the forward diffusion process, and a VAE-based encoder-decoder architecture is combined with observation-guided resampling to improve posterior sampling accuracy and diversity. The method requires no fine-tuning or additional training and is fully compatible with standard diffusion priors. Extensive experiments on ImageNet and FFHQ for super-resolution, denoising, and compressive sensing demonstrate consistent superiority over existing diffusion-based approaches—achieving higher PSNR and lower LPIPS scores—thereby validating its efficiency and strong generalization capability.
Existing DiT frameworks suffer from outdated VAE encoder backbones, low-dimensional latent spaces, and weak representational capacity due to reconstruction-based training—limiting both generation quality and training efficiency. To address these issues, we propose the Representation Autoencoder (RAE), which replaces the VAE encoder with pretrained vision representation models (e.g., DINO, SigLIP, MAE) to construct a semantically rich, high-dimensional, and scalable latent space. The decoder remains learnable and is paired with a lightweight wide DiT head, enabling efficient DiT operation in high-dimensional latent space without requiring auxiliary alignment losses for accelerated convergence. On ImageNet, RAE-DiT achieves 1.51 FID (256×256, classifier-free guidance) and 1.13 FID (256×256/512×512, with guidance), significantly improving both generative performance and training efficiency.
This work addresses the inherent uncertainty in 3D scene reconstruction from limited observations—such as a single view, sparse pixels, or noisy images—by proposing a probabilistic framework that integrates Neural Radiance Fields (NeRF) with score-based diffusion models. The method represents the 3D scene as a stochastic latent variable, employs NeRF to model the likelihood of observations, and leverages a diffusion model to learn the prior distribution over the latent space. Crucially, it introduces, for the first time, a score-based diffusion mechanism to sample from the posterior distribution of the latent variables, enabling a unified treatment of uncertainty across diverse observation conditions. A two-stage training strategy jointly optimizes the reconstruction and prior networks, achieving high-fidelity 3D reconstructions under various settings—including single-view, multi-view, noisy images, sparse pixels, and depth inputs—while faithfully capturing task-specific uncertainties.
Diffusion models on high-dimensional data suffer from poor generation quality, low training efficiency, and—critically—failure to preserve the intrinsic geometric structure of the data distribution. Method: This paper proposes a geometry-preserving encoder-decoder framework to replace conventional VAEs, enabling efficient and stable diffusion modeling in the latent space. Contribution/Results: We introduce the first theoretical encoder-decoder framework with rigorous differential-geometric constraints (e.g., isometry or conformality); provide a convergence proof for the encoder and demonstrate its acceleration effect on decoder convergence; and design a theory-guided encoder optimization strategy. Experiments show that our method significantly improves joint training stability and convergence speed, achieves superior generation quality across multiple benchmarks, and reduces training time.
This work addresses two key challenges in feed-forward native 3D generation: (1) misalignment between latent spaces and 3D geometry, and (2) the trade-off between geometric detail fidelity and computational efficiency. We propose Atlas Gaussians—a novel 3D representation that models shapes as a collection of locally UV-parameterized atlas patches, each decoded by a learnable network into a theoretically infinite 3D Gaussian point cloud. Our method introduces a patch-wise, UV-driven infinite point cloud generation paradigm, integrating local geometric-aware encoding, Transformer-based sequence modeling, and an efficient Gaussian decoding network, all trained end-to-end within a unified VAE-LDM framework. On native 3D generation tasks, our approach significantly outperforms state-of-the-art methods, producing outputs with rich geometric detail and high visual fidelity, while enabling real-time feed-forward inference.
Existing representation-based encoder generative paradigms face two key challenges: (1) discriminative feature spaces lack compact regularization, causing diffusion sampling to deviate from the data manifold and yield structural distortions; and (2) encoders exhibit weak pixel-level reconstruction capability, limiting geometric and textural fidelity. To address these, we propose a semantic-pixel joint reconstruction objective, achieving—within a compact 16×16, 96-dimensional latent space—the first unified high-semantic and high-fidelity pixel reconstruction. Our method integrates dual reconstruction losses, a compact latent-space design, a unified representation-encoder-based T2I and editing diffusion architecture, and a VAE feature-space adaptation mechanism. Experiments demonstrate significant improvements in reconstruction quality, text-to-image generation, and image editing—achieving state-of-the-art performance—along with accelerated convergence. This validates the feasibility of efficiently transferring understanding-oriented encoders into robust, generative latent spaces.
This work addresses high-fidelity 3D shape completion from a single noisy depth image. We systematically compare denoising diffusion probabilistic models (DDPMs) and autoregressive causal Transformers for generative modeling in this setting. Our key insight is that latent-space discreteness critically governs model performance: DDPMs excel at multimodal completion in continuous latent spaces, whereas autoregressive Transformers match or surpass DDPMs when operating in a unified discrete latent space. Through a discriminative baseline and rigorous ablation studies, we provide the first empirical evidence that the superiority of a generative paradigm depends more on latent-space design than on architectural choice per se. On real-world single-depth-image completion, our approach achieves state-of-the-art performance. This work establishes theoretical and practical guidance for architecture selection and latent-space design in 3D generative modeling.
This work investigates how to construct a diffusion-friendly latent space to enhance generation quality, moving beyond the sole optimization of reconstruction fidelity. The authors systematically evaluate diverse visual tokenizer architectures, regularization strategies, and latent configurations across multiple diffusion backbones. They introduce a novel metric, Velocity Irreducible Variance (VIV), to quantify velocity ambiguity in the latent space arising from trajectory intersections. Experimental results demonstrate that VIV serves as a robust predictor of generation quality, consistently outperforming other latent-space attributes across various settings. The study further uncovers several key characteristics of latent representations that exhibit strong generalization capabilities, offering actionable insights for designing better latent spaces tailored to diffusion models.
This work addresses the limitation of existing diffusion models in high-dimensional generation, which often ignore the intrinsic manifold geometry of data, while conventional latent diffusion models impose an Euclidean structure that struggles to capture complex geometries under data sparsity. To overcome this, we propose the Intrinsic Latent Diffusion Model (ILDM), the first framework to integrate Riemannian manifold geometry into the diffusion process. ILDM treats the latent space as a coordinate chart of an unknown manifold and jointly models geometric structure and uncertainty via a probabilistic decoder. We introduce a Riemannian–Euclidean hybrid forward diffusion mechanism, supported by a local uncertainty–driven diffusion strategy, a probabilistic metric tensor, and a tailored approximate denoising score matching objective. Experiments on COIL-100, MNIST, and cardiac MRI demonstrate that ILDM significantly outperforms existing methods, achieving lower FID and LPIPS scores and superior generation quality.
Diffusion models aim to construct invertible generative paths from a noise prior to the data distribution. This paper proposes a unified tripartite framework—integrating variational inference, score-based modeling, and flow matching—to reveal their shared underlying principle: continuous generative trajectories governed by time-dependent velocity fields. By formulating both the forward noising and reverse denoising processes as ordinary differential equations (ODEs), we establish a rigorous, computationally tractable continuous-time generative theory. The framework enables direct pointwise mapping at arbitrary times, flexible conditional generation with explicit control, and seamless integration with energy-based models and time-dependent neural architectures. Experimental and theoretical analyses demonstrate that this unification substantially improves sampling efficiency, controllability, and model interpretability. Our work provides a foundational theoretical framework for deepening the understanding of diffusion models and guiding the design of novel architectures.