Score
Design and implement diffusion-model–based pipelines that synthesize native 3D mesh geometry and connectivity (vertices, faces) and optionally associated texture or appearance channels; this includes building training/sampling procedures that support high face counts and resolution while preserving manifoldness and connectivity. Analyze and optimize model robustness, sampling strategies, and representation choices to improve fidelity and reliability compared with autoregressive mesh generation approaches.
This survey addresses key challenges in 3D vision—occlusion robustness, point cloud sparsity, density imbalance, and high-dimensional computational bottlenecks—across four core tasks: 3D generation, point cloud reconstruction, shape completion, and scene synthesis. Methodologically, it introduces the first unified taxonomy capturing paradigm evolution, integrating denoising diffusion probabilistic models (DDPMs), 3D conditional encoders, multi-view feature alignment, implicit neural representations (INRs), and multimodal (text/image) guidance. The work rigorously delineates current performance limits and standardizes evaluation benchmarks. Crucially, it identifies three viable technical pathways forward: efficient sampling strategies, lightweight backward processes, and large-scale 3D pretraining. These contributions provide both theoretical foundations and practical guidelines for advancing diffusion-based 3D modeling.
This work addresses the poor geometric quality in existing text-to-3D face generation methods, which often stems from irregular vertex distributions. To overcome this limitation, the authors propose constraining 3D facial geometry to a topological sphere, enabling a regular spherical representation that can be unfolded into a 2D chart. This formulation facilitates the integration of a conditional diffusion model for joint generation of geometry and texture. The approach represents the first seamless fusion of 2D diffusion models with 3D face geometry synthesis, leveraging spherical parameterization to ensure uniform point distribution, support robust mesh reconstruction, and enable geometry-guided texture synthesis. Experiments demonstrate that the method significantly outperforms current state-of-the-art techniques in text-to-3D generation, face reconstruction, and text-driven editing tasks, achieving superior geometric fidelity, textual alignment, and inference efficiency.
This work addresses the limitations of traditional mesh generation methods, which rely on sequential or autoregressive strategies and suffer from low inference efficiency and error accumulation. The authors propose an end-to-end diffusion model framework that decouples vertex and topology generation to produce high-quality, globally consistent triangular meshes. Vertices are innovatively represented as sparse voxels organized in an octree structure, and a Spacetime Interval encoding is introduced to map arbitrary edge-face topologies into continuous vertex embeddings, enabling efficient global topology recovery. Employing a coarse-to-fine strategy for vertex generation and a separate diffusion model for topology prediction, the method significantly outperforms existing autoregressive and two-stage approaches on the Objaverse and Toys4K datasets as well as on real-world images, with user studies confirming its superior perceptual quality.
Existing autoregressive 3D mesh generation methods suffer from excessive sequence lengths and high computational costs due to flattening meshes into vertex sequences. To address this, this work proposes FACE, a novel framework that performs autoregressive modeling at the face (triangle) level using a “one-face-one-token” strategy, reducing sequence length by a factor of nine (compression ratio of 0.11) and substantially improving generation efficiency. The approach integrates a face-level autoregressive autoencoder (ARAE), a VecSet encoder, and a latent diffusion model. Evaluated on standard benchmarks, FACE achieves state-of-the-art reconstruction quality and successfully enables high-fidelity single-image-to-3D-mesh generation.
Existing AI-based 3D mesh generation suffers from low topological efficiency and uncontrollable face counts. Method: This paper introduces the first face-level triangular mesh streaming diffusion Transformer (DiT) framework. It pioneers streaming diffusion for face-level generation, designs a synergistic architecture combining a face-level VAE and a conditional diffusion Transformer, and enables non-autoregressive generation with continuous-space diffusion and precise face-count control. Contributions/Results: Our method generates 800-face meshes in just 3.2 seconds—35× faster than state-of-the-art methods—and achieves superior performance on ShapeNet and Objaverse. It supports multimodal conditioning—including text and sketch inputs—significantly reducing manual modeling effort. This work establishes a new paradigm for efficient, controllable, and high-quality 3D content generation.
This work addresses the challenge of jointly preserving global structure and local geometric detail in high-fidelity 3D shape generation. We propose a two-stage diffusion framework: the first stage generates coarse-grained voxelized global structures, while the second stage refines geometric details via spatially localized voxel queries and Rotary Position Embedding (RoPE) for precise spatial anchoring. Innovatively, we introduce watertightness-preserving preprocessing and a geometry-decoupled refinement mechanism to ensure topological integrity and surface fidelity under resource constraints. The method integrates 3D diffusion modeling, voxel-based representation, RoPE-enabled spatial localization, watertight mesh repair, and multi-level data augmentation. Trained solely on public 3D datasets, our approach achieves state-of-the-art geometric quality—significantly outperforming existing open-source methods—while supporting end-to-end high-quality generation and full open-source reproducibility.
Existing autoregressive approaches to native 3D mesh generation are limited by constraints on face count, vertex resolution, and texture support. This work proposes the Barycentric Dominance Field (BDF), which for the first time encodes discrete mesh topological connectivity as a continuous signal defined over the surface of a triangle mesh. This formulation seamlessly integrates with mainstream 3D diffusion-based generative models without requiring any architectural modifications. The method enables high-quality, scalable, and robust native mesh synthesis, significantly outperforming state-of-the-art autoregressive methods in both geometric detail fidelity and topological completeness.
This study addresses the slow generation of autoregressive models and the reliance of continuous flow models on heuristic decoders by proposing MeshOctave. This framework introduces a novel octree-based binary spatial mesh scale definition that treats coarsening as deterministic vertex merging, thereby supporting unordered set modeling and dynamic adaptive refinement. Furthermore, it employs a scale-conditioned masked uniform discrete diffusion model to learn split-reconnection operations, enabling parallel, serialization-free hierarchical mesh generation. Experimental results demonstrate that MeshOctave significantly outperforms baseline methods in both geometric fidelity and topological validity, while naturally extending to mesh subdivision tasks.
Direct generation of triangular meshes faces significant challenges in modeling the permutation symmetries of faces and their constituent vertices, and conventional autoregressive approaches suffer from low efficiency. This work proposes an equivariant optimal transport flow matching model that directly generates unordered triangle soups, thereby circumventing sequential modeling and rigorously preserving equivariance under arbitrary permutations of faces and their internal vertices. To this end, we introduce a mesh-specific equivariant Diffusion Transformer architecture and employ an optimal transport–based training objective that eliminates symmetry-breaking supervisory signals. Experimental results demonstrate that our method achieves generation quality on par with state-of-the-art autoregressive models while offering approximately 18× faster inference speed.
Existing generative models struggle to achieve large-scale geometric deformations guided by images while preserving both the 3D mesh topology and part-level semantic structure. To address this challenge, this work proposes an image-driven framework for 3D mesh geometric stylization. The method leverages a pretrained diffusion model to extract abstract geometric representations from target images, combines differentiable rendering with an approximate VAE encoder to provide stable gradients, and employs a coarse-to-fine multi-scale deformation strategy to enable diverse style transfer. This approach is the first to effectively maintain original topological and semantic integrity under significant geometric deformation, successfully generating stylized 3D meshes that faithfully reflect the pose and contour characteristics of the input images, thereby substantially enhancing the expressiveness of artistic 3D content creation.
Existing 3D generation methods rely on high-dimensional geometric representations—such as voxels, signed distance functions (SDFs), or point clouds—which incur substantial computational and memory costs, making it challenging to simultaneously achieve high resolution and controllability. This work introduces the first diffusion-based approach operating directly in the compact parameter space of superquadrics, representing 3D shapes with only approximately 7 KB of parameters encoding pose, scale, and shape. By drastically reducing the state dimensionality, the method enables resolution-free point cloud decoding, part-level editing, and explicit geometric constraints. It achieves competitive surface fidelity and distribution quality on standard benchmarks, with per-shape generation times consistently under 0.6 seconds.