ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost of high-fidelity 3D generation, which typically relies on large-scale data and models while underutilizing the rich semantic and structural priors embedded in discriminative 3D foundation models. To bridge this gap, the authors propose ROAD, a novel framework that, for the first time, transfers priors from discriminative 3D foundation models into a diffusion Transformer. ROAD introduces a reciprocal objective alignment mechanism that effectively reconciles the heterogeneity between generative and discriminative latent spaces through global semantic compression and optimal micro-structure matching—formulated as bipartite graph matching. Remarkably, without increasing inference overhead, ROAD achieves generation quality on par with the industrial baseline Step1X-3D using only 1.5% of the training data, substantially reducing both training cost and computational requirements.
📝 Abstract
High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requires learning geometry from scratch and overlooks the rich semantic and structural priors already encapsulated in discriminative 3D foundation models. We contend that leveraging the profound understanding of the 3D world possessed by these discriminative models can significantly reduce generative cost. To this end, we propose ROAD, a framework that reduces the training cost of 3D generation by transferring these rich discriminative priors into diffusion transformers. To address the inherent semantic-structural heterogeneity between generative and discriminative latents, we introduce a reciprocal-objective alignment strategy. This method synergizes Holistic Semantic Condensing to enforce global semantic coherence and Structural Optimal Alignment, which is formulated as a bipartite matching problem to rigorously align microscopic geometric details between disparate latent spaces. The 3D foundation model is only used for training-time supervision of alignment and is not used at inference, incurring no additional inference cost. Compared with the industrial baseline Step1X-3D, the proposed ROAD achieves highly competitive generation performance with only 1.5% of the training data and significantly reduces training costs, effectively reducing the computational overhead of high-fidelity 3D generation. Code is available at https://github.com/H-EmbodVis/ROAD.
Problem

Research questions and friction points this paper is trying to address.

3D shape generation
computational cost
discriminative priors
semantic-structural heterogeneity
high-fidelity generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reciprocal-Objective Alignment
Discriminative Prior Transfer
Diffusion Transformer
Semantic-Structural Heterogeneity
3D Shape Generation
🔎 Similar Papers
No similar papers found.
X
Xiao Luo
Huazhong University of Science and Technology, China
M
Mingyang Du
Huazhong University of Science and Technology, China
Xin Zhou
Xin Zhou
Huazhong University of Science and Technology
Computer Vison3D Vision
T
Tianrui Feng
Huazhong University of Science and Technology, China
X
Xiwu Chen
Megvii, China
Xiaofan Li
Xiaofan Li
East China Normal University
Computer Vision
J
Jiangning Zhang
Zhejiang University, China
Dingkang Liang
Dingkang Liang
Huazhong University of Science and Technology
Embodied AIWorld ModelAutonomous DrivingCrowd Counting