🤖 AI Summary
This study addresses the disconnection between scene reconstruction and policy development, as well as the challenges of sim-to-real transfer. We propose a unified agent framework that establishes an end-to-end closed loop from video input to robotic manipulation. By integrating a unified task interface, the framework seamlessly connects scene reconstruction, code generation, and real-world execution. It iteratively refines scene representations through depth estimation and visual feedback mechanisms, while leveraging large model-based coding agents to automatically generate observation-driven execution policies. Experimental results demonstrate that the system achieves a mean depth error of only 0.1057 meters, with real-robot task success rates reaching 80% of their simulation counterparts. These findings validate the high-fidelity transfer capability of the proposed approach.
📝 Abstract
A simulation of a real robot workspace must preserve task-relevant interactions, while policies developed in it must operate on observations available to the real robot. Yet scene reconstruction and policy development are often treated separately. We present Agentic Real-to-Sim-to-Real (Agentic RSR), a framework that links scene reconstruction, policy development, and real-robot execution through the same manipulation task. Given a workspace video, a task description, and a known robot model, an agent recovers metric scale, iteratively refines the scene using visual feedback, and checks task-relevant interactions in MuJoCo. A coding agent then develops an executable policy, progressing from privileged object poses to visual observations and randomized simulation. The policy can interleave multiple observations and actions within one invocation, while the agent uses execution feedback to continue, retry, or revise its approach. A shared task-level interface carries the policy and accumulated experience to the real robot, where fresh observations and safety checks guide execution. Across 18 reconstructed scenes involving two robots, the mean four-view Depth MAE against reference depth estimates is 0.1057 m, the mean Lab $ΔE_{76}$ is 11.04, and the mean grayscale SSIM is 0.6990. In real-robot experiments, the aggregate task success rate reaches 80% of the simulation task success rate, indicating substantial retention of simulated performance on hardware. Code and reconstructed scene data will be made publicly available.