Learning Social Navigation from Internet Videos in the Policy State Space
This study addresses the challenge that training social navigation policies relies on expensive simulators and predefined pedestrian behaviors. To overcome this, we propose a framework that directly transforms ordinary monocular videos into closed-loop training environments. The core innovation lies in integrating metric traversability maps with pedestrian trajectory replay to define a dynamics model within the policy state space, enabling efficient simulation of counterfactual robot states without 3D reconstruction or photorealistic rendering. Experimental results demonstrate that our method achieves an 81.2% success rate on the Arena benchmark, outperforming the strongest baseline. Furthermore, real-world deployment attains 19 out of 20 successful trials without fine-tuning, validating its strong generalization capability.