🤖 AI Summary
This study addresses the limitation of existing single-image 3D generation methods in producing physics-simulation-ready scenes, which commonly suffer from gravity misalignment and contact instability. We propose a grounding refinement framework that achieves asset decoupling through instance- and ground-level prompting, optimizes object poses via differentiable rendering, and introduces physical constraints to rectify contact relationships. This approach effectively transforms visual assets into simulation-ready representations capable of supporting agent perception-reasoning-action loops. Experimental results demonstrate that the proposed method significantly improves viewpoint alignment accuracy, contact plausibility, and physical stability, thereby facilitating goal-directed interactive tasks.
📝 Abstract
Agentic recognition requires visual perception to move beyond static scene understanding and produce structured scene representations that support the perception--reasoning--action loop. Existing single-image 3D generation methods, however, mainly produce visually plausible object assets rather than simulation-ready scene states. When independently generated meshes are composed in a shared space, they may fail to align with the input camera, violate gravity, interpenetrate nearby objects, or become unstable under physics simulation. We present Fysiverse-3D-SimReady, a grounded refinement framework for reconstructing simulation-ready multi-object scenes from a single RGB image with instance and ground prompts. The method places generated object meshes into a shared gravity-aligned scene, refines their camera-space poses through differentiable rendering, and corrects scene-level supports and contacts for stable physical execution. The scene is then used by an agentic simulation workflow, which converts a scene-specific task goal into an executable physics rollout rendered from the original camera view. Experiments show that Fysiverse-3D-SimReady improves input-view alignment, contact plausibility, and physical stability over existing single-image reconstruction and scene generation baselines, while enabling goal-conditioned physical interactions from a single image.