D3GS: Depth, DINO, and RGB Diffusion Co-Guided 3D Gaussian Splatting for Sparse-View Reconstruction
为解决3D高斯点云从稀疏视角重建时的几何模糊、视图不一致和细节缺失问题,提出D3GS框架,通过深度-迪诺-扩散指导方法联合增强几何与外观。
为解决3D高斯点云从稀疏视角重建时的几何模糊、视图不一致和细节缺失问题,提出D3GS框架,通过深度-迪诺-扩散指导方法联合增强几何与外观。
This work addresses the limitations of existing methods for generating simulation-ready 3D assets from a single image, which suffer from implicit reasoning that entangles part layout with local shape and lacks supervisability over intermediate states. To overcome this, we propose an explicit, structured physical reasoning framework that models part decomposition, 2D/3D localization, inter-part relationships, coarse geometry, and surface cues sequentially through interpretable state trajectories, enabling supervision, conditional control, and optimization of intermediate steps. Our approach employs factorized decoding—using 3D bounding boxes for pose and local codes for shape—alongside a Chain-of-Thought-aligned GRPO algorithm and a frozen decoder architecture. Evaluated under a unified protocol, our method consistently outperforms baselines across geometric, scale, and physical plausibility metrics, producing assets that exhibit high-fidelity parseability, accurate collision responses, and functional articulation in Unreal Engine 5.
为解决3D高斯点云从稀疏视角重建时的几何模糊、视图不一致和细节缺失问题,提出D3GS框架,通过深度-迪诺-扩散指导方法联合增强几何与外观。
This work addresses the limitations of existing methods for generating simulation-ready 3D assets from a single image, which suffer from implicit reasoning that entangles part layout with local shape and lacks supervisability over intermediate states. To overcome this, we propose an explicit, structured physical reasoning framework that models part decomposition, 2D/3D localization, inter-part relationships, coarse geometry, and surface cues sequentially through interpretable state trajectories, enabling supervision, conditional control, and optimization of intermediate steps. Our approach employs factorized decoding—using 3D bounding boxes for pose and local codes for shape—alongside a Chain-of-Thought-aligned GRPO algorithm and a frozen decoder architecture. Evaluated under a unified protocol, our method consistently outperforms baselines across geometric, scale, and physical plausibility metrics, producing assets that exhibit high-fidelity parseability, accurate collision responses, and functional articulation in Unreal Engine 5.