Score
Designs and implements convolutional generator architectures and training pipelines that convert 2D image inputs into 3D outputs (meshes, voxel grids, point clouds, or textured volumes). Engineers convolutional application across large images or tiled inputs to preserve local detail, maintain spatial consistency, and synthesize coherent 3D assets.
This paper presents a systematic survey of deep learning–driven 3D shape generation, addressing three core dimensions: shape representation, generative modeling, and evaluation protocols. Methodologically, it introduces the first unified taxonomy covering explicit (e.g., meshes), implicit (e.g., SDFs, NeRFs), and hybrid representations, traces the evolution of feedforward-based generative architectures, and consolidates major benchmarks (e.g., ShapeNet, FAUST) and metrics (e.g., Chamfer distance, Jensen–Shannon divergence), revealing inherent trade-offs among fidelity, diversity, and realism. Its principal contribution is a novel “representation–model–evaluation” triadic analytical framework, which explicitly identifies controllable shape modeling, efficient inference, and physically consistent generation as key open challenges. The framework establishes a structured benchmark and roadmap for future research in 3D generative modeling. (126 words)
Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.
3D generative models face fundamental trade-offs between geometric fidelity, computational efficiency, and scalability due to the difficulty of modeling complex spatial structures and the high cost of volumetric representation. To address this, we propose VoxSet—a semi-structured voxel representation that explicitly encodes spatial topology under high compression, enabling position-aware generation and token-level test-time scaling. Our two-stage framework first generates sparse voxel anchors, then refines geometry to arbitrary resolution via a correction-flow Transformer. This unifies structured 3D modeling with efficient decoding, drastically reducing training overhead. Evaluated across multiple benchmarks, VoxSet achieves state-of-the-art performance in fidelity, diversity, and scalability metrics. It enables high-fidelity, large-scale, and flexible inference for 3D asset generation—supporting resolution-agnostic output and interactive editing without retraining.
Current single-image 3D generation methods suffer from insufficient geometric detail, over-smoothed surfaces, and structural discontinuities—particularly in thin-shell geometries—rendering them inadequate for industrial-grade applications. To address these limitations, we propose a multi-dimensional collaborative optimization framework: (1) a geometry-aware implicit 3D representation tailored for high-fidelity surface modeling; (2) a linear Transformer architecture to enhance long-range geometric coherence; and (3) a progressive super-resolution strategy integrated with strengthened multi-view 3D supervision. Our approach significantly improves geometric fidelity and structural integrity, achieving state-of-the-art performance on ShapeNet and Objaverse benchmarks. The generated models exhibit high-precision geometry, topologically consistent thin-wall structures, and immediate usability—enabling seamless integration into professional 3D production pipelines.
To address the remeshing bottleneck in triangular mesh 3D model classification caused by topological irregularity, this paper proposes a native convolutional neural network architecture specifically designed for triangle meshes. Our method introduces two key innovations: (1) a face-based local convolution operator grounded in face adjacency relationships, explicitly modeling geometric and topological dependencies among mesh faces; and (2) a differentiable, learnable face-collapsing pooling mechanism that enables hierarchical feature dimensionality reduction while preserving topological awareness. The architecture operates directly on raw triangular meshes—eliminating the need for remeshing—and is fully end-to-end trainable. Evaluated on ModelNet40, ShapeNet Part, and FAUST for semantic classification, our approach achieves state-of-the-art or competitive accuracy, while significantly reducing memory consumption (−37% on average) and computational cost (−41% FLOPs), effectively balancing representational power and efficiency.
This work addresses the end-to-end reconstruction of complex 3D assets from a single RGB image. We propose a neural procedural graph generation framework that represents 3D structure via differentiable procedural graphs, introduces an edge-based tokenization strategy, leverages Transformers to model structural sequence priors, and—crucially—incorporates Monte Carlo Tree Search (MCTS) for guided sampling, significantly improving image-to-3D alignment accuracy. Unlike prior approaches, our method requires neither category-specific priors nor multi-view supervision, enabling direct generation of decodable and editable 3D assets from monocular images. Evaluated on diverse complex objects—including cacti, trees, and bridges—our approach outperforms existing generative 3D methods and domain-specific modeling techniques in both fidelity and generalizability. Notably, it demonstrates strong generalization to real-world images while preserving fine-grained geometric and topological structure.
To address the demand for high-fidelity, diverse 3D asset generation and flexible editing, this paper introduces SLAT—a structured 3D implicit representation that jointly encodes sparse 3D mesh topology and multi-view visual foundation model features, enabling unified decoding into multiple 3D formats (e.g., radiance fields, 3D Gaussians, explicit meshes). Methodologically, SLAT pioneers a 2B-parameter Transformer architecture based on Rectified Flow for large-scale 3D latent-space modeling—the first of its kind. We curate a high-quality dataset of 500K 3D assets and perform end-to-end training. SLAT supports text- and image-conditioned generation, achieving state-of-the-art performance in fidelity, diversity, and editability. It enables real-time local 3D editing and dynamic output format switching. All code, models, and data are publicly released.
This study addresses the lack of standardization and poor reproducibility in data generation for machine learning modeling of three-dimensional obstructed channel flows. We propose a configuration-driven, end-to-end automated framework integrating parametric CAD modeling, signed distance field (SDF)-based voxelization, high-fidelity lattice Boltzmann simulations using waLBerla, and multi-resolution tensor-based registration—all orchestrated via Hydra/OmegaConf to enable fully configurable pipelines and systematic ablation studies. Our key contributions are: (1) the first standardized data generation paradigm specifically designed for obstructed flows, supporting joint geometric–flow-field parameterization; and (2) a large-scale, high-quality 3D flow dataset comprising over 10,000 samples spanning Reynolds numbers Re = 100–15,000. The dataset demonstrates superior storage efficiency and empirical effectiveness in training physics-informed models (e.g., 3D U-Net), significantly enhancing reproducibility and generalizability in physics-guided machine learning.
Existing 3D generative models often struggle to achieve pixel-level fidelity when synthesizing 3D assets from images due to ambiguities in 2D–3D correspondences. This work proposes a pixel-aligned 3D generation paradigm that explicitly lifts multi-scale 2D image features into a 3D feature volume consistent with the input viewpoint via a pixel reprojection mechanism, integrated within an end-to-end trainable, native 3D generative model. The approach achieves, for the first time, large-scale, natively 3D pixel-aligned synthesis, significantly improving geometric and appearance fidelity under both single-image and multi-view settings—approaching the quality of traditional reconstruction methods—and successfully extends to high-fidelity, object-disentangled scene-level synthesis tasks.
Existing methods struggle to generate large-scale 3D scenes that simultaneously maintain global structural consistency and support fine-grained layout control. This work presents the first extension of single-image 3D generation models to the scene scale, introducing a convolutional 3D diffusion architecture and a dedicated synthetic data engine to mitigate the scarcity of real-world 3D scene datasets. By reformulating the image-to-3D generator as a convolutional operator and fine-tuning it on synthetically generated scenes, the proposed approach enables the generation of arbitrarily sized and complex 3D environments that are both geometrically detailed and structurally coherent. The method significantly outperforms existing approaches across diverse layouts and semantic prompts, demonstrating strong controllability and scalability in 3D scene synthesis.
Converting analog circuit reference layouts into programmable layout generators remains inefficient and labor-intensive. Method: This paper proposes a convolutional neural network (CNN)-based automated translation method that performs end-to-end learning to accurately classify layout subcell instances and automatically match them to an existing generator library, recommending their optimal hierarchical placement. Contribution/Results: To the best of our knowledge, this is the first work to apply CNNs to the “layout-to-generator” mapping task, significantly enhancing generalization—enabling accurate identification of unseen subcells with substantial structural divergence from training samples. Evaluated on a dataset of 4,885 instances, the method achieves 99.3% classification accuracy. Processing time per instance drops from 88 minutes manually to 18 seconds automatically, markedly improving layout reuse and generator-based layout synthesis efficiency for analog circuits.
This work addresses the scarcity of high-quality, structured training data that hinders 3D content generation by introducing an open-source 3D asset ecosystem. It proposes a hybrid data strategy that integrates real-world, high-fidelity 3D objects with AI-generated assets covering long-tail categories, enriched with part-level semantic annotations to enable fine-grained editing and perception. Leveraging high-fidelity mesh processing, multi-view rendering, and scalable AIGC-based synthesis techniques, the project releases a large-scale dataset comprising 250,000 real and 125,000 synthetic 3D assets. This dataset effectively supports the training of the Hunyuan3D-2.1-Small model, significantly advancing the application of 3D generative models across multiple domains.