Score
Designs and builds complete 3D assets — geometry, materials, textures, rigs, and supporting metadata — packaged for use in production pipelines. Specifically composes and authors assets in the USD (Universal Scene Description) format by creating USD primitives, layers, references, variant sets, and stage metadata to produce reusable USD asset packages.
This work proposes the first autoregressive Transformer-based text-to-3D generation framework tailored for user-generated content (UGC) scenarios, where there is a growing demand for high-quality, diverse, and design-constrained modular 3D assets. The method represents modular 3D assets as sequences of primitive elements and leverages sequence modeling and decoding mechanisms inspired by language models to generate structurally coherent and semantically consistent assets. By introducing a novel modular sequential representation and an autoregressive decoding strategy, the framework establishes a scalable and general-purpose generative architecture. Extensive evaluation on a real-world dataset collected from an online platform demonstrates that the proposed approach significantly improves generation quality, making it well-suited for both professional development and UGC applications.
Existing 3D generation methods struggle to meet production-grade requirements for real-time interactive applications, such as consistent topology, UV unwrapping, physically based rendering (PBR) materials, skeletal rigging, and physically plausible scene layout. To address this gap, this work proposes a two-dimensional taxonomy centered on asset production pipelines—structured by asset type and production stage—and systematically constructs a comprehensive generation framework encompassing geometry synthesis, topology optimization, UV parameterization, PBR appearance modeling, skeletal rigging, and physics-aware scene assembly. The authors further introduce a cross-dimensional evaluation protocol to rigorously assess the direct usability of generated assets in game engines and simulation platforms. Their analysis highlights critical challenges in data quality, controllable generation, and end-to-end assetization, underscoring the pivotal role of deployable 3D content as foundational infrastructure for embodied intelligence and interactive world models.
To address the demand for high-fidelity, diverse 3D asset generation and flexible editing, this paper introduces SLAT—a structured 3D implicit representation that jointly encodes sparse 3D mesh topology and multi-view visual foundation model features, enabling unified decoding into multiple 3D formats (e.g., radiance fields, 3D Gaussians, explicit meshes). Methodologically, SLAT pioneers a 2B-parameter Transformer architecture based on Rectified Flow for large-scale 3D latent-space modeling—the first of its kind. We curate a high-quality dataset of 500K 3D assets and perform end-to-end training. SLAT supports text- and image-conditioned generation, achieving state-of-the-art performance in fidelity, diversity, and editability. It enables real-time local 3D editing and dynamic output format switching. All code, models, and data are publicly released.
Single-image 3D scene generation with multiple objects faces severe challenges including heavy occlusion and object coupling-induced geometric distortion. To address these, we propose a two-stage differentiable framework: first, leveraging off-the-shelf image-to-3D models to independently reconstruct per-object meshes; second, jointly optimizing global scene layout via differentiable rendering, incorporating a novel optimal transport-driven long-range appearance loss and a high-level semantic loss in a synergistic constraint mechanism—enabling unified modeling of object-level geometric independence and scene-level structural consistency. Our approach integrates differentiable rendering, optimal transport theory, and semantics-guided gradient optimization. Evaluated on multi-object benchmarks, our method significantly improves geometric detail fidelity, object separation, and global coherence, outperforming state-of-the-art single-image 3D generation methods both quantitatively and qualitatively.
To address the challenge of jointly achieving precise layout control and faithful attribute rendering in multi-instance text-to-image generation (MIG), this paper proposes a two-stage decoupled framework. In the first stage, LDM3D generates high-fidelity depth maps to enable fine-grained instance localization and holistic scene structure modeling. In the second stage, a pre-trained ControlNet performs zero-shot conditional rendering, enabling plug-and-play integration with arbitrary base diffusion models (e.g., SD2, SDXL) without fine-tuning. We introduce a novel depth-driven compositional paradigm and design a lightweight depth adapter to enhance layout controllability. Evaluated on COCO-Position and COCO-MIG benchmarks, our method achieves significant improvements in layout accuracy and attribute fidelity while demonstrating strong generalization and model-agnostic compatibility. The implementation is publicly available.
This work proposes the first end-to-end method for generating high-quality, editable 3D assets from a single image of indoor furniture or decor, tailored to the demands of interior design and e-commerce applications. The system employs a modular architecture comprising four coordinated stages—geometry reconstruction, texture generation, material assignment, and part decomposition—to produce watertight meshes with physically based rendering (PBR) materials and semantic part labels. Key innovations include implicit signed distance field (SDF) modeling via a geometry VAE and DiT, multi-view back-projection coupled with 3D texture field completion, MatWeaver-driven material matching, and multi-part joint SDF decoding enabled by PartVAE and PartDiT. Experiments demonstrate state-of-the-art performance across dedicated metrics, yielding high-fidelity 3D assets with accurate materials, structural completeness, and semantic editability.
This work addresses the limitations of existing 3D generation methods, which typically produce static, opaque meshes devoid of semantic structure and procedural control, thereby hindering interactive editing. The authors propose the first end-to-end, code-native framework for 3D asset generation that leverages large language models to directly synthesize executable Blender Python scripts from multimodal inputs—text and images—yielding glTF assets enriched with named components, hierarchical assemblies, constraints, and articulated joints. Evaluated on Nova3D-Bench, a newly introduced benchmark comprising 54 diverse tasks, the method consistently generates valid programs and high-fidelity geometric assets, satisfies over 98% of numerical constraints, achieves 100% fidelity in local edits, and successfully constructs 59 functional joints, substantially outperforming current baselines.
Existing 3D generation methods struggle to simultaneously achieve high visual fidelity, real-time performance, and mobile deployment. This work proposes the first single-image 3D generation framework that balances deployment efficiency and interactive speed, producing high-quality meshes with baked normals, colored textures, and controllable face counts within 30 seconds; its Flash variant delivers preview-quality results in just 14 seconds. The approach integrates coarse-to-fine VecSet-based geometry generation, multi-view texture synthesis, and 3D back-projection inpainting, while performing mesh simplification, cleanup, normal baking, and parallel UV unwrapping directly on the GPU. Combined with model distillation and pipeline parallelism, the system minimizes end-to-end latency. Experiments demonstrate that the generated assets match the visual quality of commercial solutions, with both automated metrics and blind human evaluations confirming the method’s efficiency and practicality.
This study addresses the challenge faced by production system engineers in automatically verifying production line layouts due to limited knowledge of PDDL and planning theory. To bridge this gap, the authors propose a novel approach based on an Asset Administration Shell (AAS) capability model that natively generates complete PDDL planning problems directly from domain-level descriptions, eliminating the need for PDDL-specific submodels. The method integrates four Industry 4.0 standards—VDI 3682, IEC 61360-1, IDTA 02011, and IDTA 02016—to construct the AAS and employs an extraction algorithm to automatically translate multi-AAS architectures into PDDL domains. In a laboratory case study, the approach enabled engineers to systematically compare four layout variants by modifying only the AAS model, significantly lowering the barrier to adopting automated planning in industrial settings.