Score
Designs and implements generative diffusion models and denoising processes that produce temporally evolving 3D mesh vertex positions (mesh trajectories), operating directly in world coordinates to sample diverse, plausible future motions and capture uncertainty in dynamics. Builds, evaluates, and analyzes diffusion-based trajectory representations and samplers over sequences of mesh states, ensuring temporal consistency and realistic geometric motion.
This survey addresses key challenges in 3D vision—occlusion robustness, point cloud sparsity, density imbalance, and high-dimensional computational bottlenecks—across four core tasks: 3D generation, point cloud reconstruction, shape completion, and scene synthesis. Methodologically, it introduces the first unified taxonomy capturing paradigm evolution, integrating denoising diffusion probabilistic models (DDPMs), 3D conditional encoders, multi-view feature alignment, implicit neural representations (INRs), and multimodal (text/image) guidance. The work rigorously delineates current performance limits and standardizes evaluation benchmarks. Crucially, it identifies three viable technical pathways forward: efficient sampling strategies, lightweight backward processes, and large-scale 3D pretraining. These contributions provide both theoretical foundations and practical guidelines for advancing diffusion-based 3D modeling.
This work addresses the black-box nature of the denoising process in Denoising Diffusion Probabilistic Models (DDPMs) for 2D point cloud generation. We propose InJecteD, the first framework to systematically quantify trajectory dynamics—including displacement, velocity, clustering structure, and drift field evolution—during denoising. By simplifying the DDPM architecture and integrating Fourier time embeddings with adaptive noise scheduling, we analyze trajectory evolution across four configuration variants using Wasserstein distance and cosine similarity. Experiments uncover dataset-specific denoising mechanisms: e.g., bullseye exhibits concentric convergence, while dino demonstrates progressive contour formation. Fourier time embeddings significantly improve trajectory stability and reconstruction fidelity. Our approach enhances the interpretability of diffusion models by establishing a quantitative, analyzable paradigm for human-in-the-loop debugging and generative process control.
This work addresses the limitations of traditional mesh generation methods, which rely on sequential or autoregressive strategies and suffer from low inference efficiency and error accumulation. The authors propose an end-to-end diffusion model framework that decouples vertex and topology generation to produce high-quality, globally consistent triangular meshes. Vertices are innovatively represented as sparse voxels organized in an octree structure, and a Spacetime Interval encoding is introduced to map arbitrary edge-face topologies into continuous vertex embeddings, enabling efficient global topology recovery. Employing a coarse-to-fine strategy for vertex generation and a separate diffusion model for topology prediction, the method significantly outperforms existing autoregressive and two-stage approaches on the Objaverse and Toys4K datasets as well as on real-world images, with user studies confirming its superior perceptual quality.
Generating continuous textures (e.g., RGB images) directly on 3D mesh surfaces remains challenging, as existing approaches rely on mesh parameterization or implicit representations—both prone to geometric distortion and fidelity loss. This paper introduces the first probabilistic generative framework that jointly leverages heat diffusion (governed by the Laplacian–Beltrami operator) and denoising diffusion to model texture signals on non-Euclidean manifolds while preserving intrinsic geometric structure. Unlike prior methods, it operates directly on mesh topology without parameterization or implicit fields, enabling topology-aware feature propagation and shape-conditioned texture synthesis. Experiments demonstrate high-fidelity, geometry-consistent RGB texture generation on complex surfaces. Notably, our method achieves, for the first time, category-level conditional texture synthesis across diverse shapes—marking a paradigm shift in 3D asset generation.
This work uncovers a universal low-dimensional geometric regularity in deterministic sampling trajectories of diffusion models: all trajectories lie exactly within an extremely low-dimensional linear subspace and consistently exhibit a “boomerang”-shaped structure—irrespective of model architecture, conditioning inputs, or generated content. To characterize and exploit this phenomenon, we first formally model the geometric structure of sampling trajectories and propose a trajectory analysis framework grounded in probability flow ODEs and kernel density estimation. Building upon this, we design a dynamic programming–driven schedule alignment strategy that jointly improves sampling efficiency and quality using only 5–10 function evaluations. Our method is lightweight, requires no retraining, and is fully compatible with mainstream ODE solvers (e.g., DOPRI5), incurring negligible computational overhead.
Deformable object manipulation—particularly for cloth—faces challenges in state estimation and dynamics modeling due to high dimensionality, strong nonlinearity, and frequent self-occlusion. To address these, this paper introduces the first Transformer-based diffusion model framework specifically designed for deformable objects. Our method jointly achieves full-state reconstruction from sparse RGB-D observations and action-conditioned long-horizon dynamics prediction. It pioneers the integration of diffusion generative modeling into the perception–control closed loop for deformable objects, decoupling and co-optimizing perception and dynamics modeling to overcome the local-receptive-field limitations of graph neural networks (GNNs). This enables global deformation representation and high-fidelity state generation. Experiments demonstrate significantly improved state reconstruction accuracy and a tenfold reduction in long-horizon prediction error. Furthermore, our approach successfully executes multi-step cloth folding tasks on a real robotic platform.
Existing 3D generation methods rely on high-dimensional geometric representations—such as voxels, signed distance functions (SDFs), or point clouds—which incur substantial computational and memory costs, making it challenging to simultaneously achieve high resolution and controllability. This work introduces the first diffusion-based approach operating directly in the compact parameter space of superquadrics, representing 3D shapes with only approximately 7 KB of parameters encoding pose, scale, and shape. By drastically reducing the state dimensionality, the method enables resolution-free point cloud decoding, part-level editing, and explicit geometric constraints. It achieves competitive surface fidelity and distribution quality on standard benchmarks, with per-shape generation times consistently under 0.6 seconds.
This study addresses the problem of constructing intrinsic diffusion processes on unknown low-dimensional manifolds using only point cloud data, without access to explicit geometric information such as coordinate charts or projections. To this end, the authors propose the Implicit Manifold Diffusion (IMD) framework, which estimates the infinitesimal generator and carré-du-champ operator via a neighborhood graph to formulate a data-driven stochastic differential equation in the ambient space whose dynamics faithfully encode the intrinsic geometry of the underlying manifold. Theoretical analysis establishes that the induced probability paths weakly converge to the ideal diffusion process on the true manifold as the sample size grows. Numerically, the method is efficiently implemented using Euler–Maruyama integration. This work presents the first provably convergent manifold diffusion model in a fully implicit setting, offering both theoretical guarantees and practical tools for manifold-aware generation and sampling.
This work addresses the challenge of generating physically plausible and geometry-aware 3D motion simulations for multiple objects in world coordinates, accommodating both rigid and elastic materials as well as complex interactions. To this end, the authors propose PhysiFormer, the first diffusion-based Transformer model that directly predicts future trajectories of 3D mesh vertices in world space, eschewing traditional physics constraints and view-dependent representations. PhysiFormer employs a temporally-, spatially-, and object-wise factorized attention mechanism to achieve permutation-invariant multi-object reasoning. Taking vertex positions, velocities, and material types as input, the model leverages a denoising diffusion process to produce diverse yet physically consistent motion sequences. Experiments demonstrate that PhysiFormer significantly outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum consistency, and generalizes effectively to scenarios involving mixed materials, real-world geometries, and larger numbers of interacting objects.
This work addresses the limitation of existing diffusion models in generating incompressible flow fields, which often neglect physical constraints or enforce them only through soft penalties. To overcome this, we propose a diffusion-based generative framework that integrates hard geometric constraints with soft physical regularization. Our approach combines boundary-condition-guided diffusion, a physics-informed loss incorporating divergence penalties, and a projection-constrained reverse sampling scheme based on a geometry-aware Helmholtz–Hodge decomposition. Furthermore, we bridge the generative model with the intrinsic geometry of incompressible flows via constrained Langevin dynamics on manifolds. Experiments demonstrate that our method significantly outperforms baseline approaches on both analytical Navier–Stokes solutions and complex obstacle scenarios, achieving substantial improvements in divergence error, spectral accuracy, vorticity statistics, and boundary consistency.