Learning Social Navigation from Internet Videos in the Policy State Space

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that training social navigation policies relies on expensive simulators and predefined pedestrian behaviors. To overcome this, we propose a framework that directly transforms ordinary monocular videos into closed-loop training environments. The core innovation lies in integrating metric traversability maps with pedestrian trajectory replay to define a dynamics model within the policy state space, enabling efficient simulation of counterfactual robot states without 3D reconstruction or photorealistic rendering. Experimental results demonstrate that our method achieves an 81.2% success rate on the Arena benchmark, outperforming the strongest baseline. Furthermore, real-world deployment attains 19 out of 20 successful trials without fine-tuning, validating its strong generalization capability.
📝 Abstract
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Social Navigation
Simulator Construction
Pedestrian Behavior
Training Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Social Navigation
Internet Videos
Policy State Space
Traversability Map
Counterfactual Simulation
🔎 Similar Papers
Jiaming Wang
Jiaming Wang
National University of Singapore
Generative AIRobotics
D
Duc Thang Nguyen
National University of Singapore, Singapore
J
Jizhuo Chen
National University of Singapore, Singapore
V
Volodymyr Shcherbyna
Singapore Management University, Singapore
D
Diwen Liu
The Chinese University of Hong Kong, Hong Kong SAR, China
Z
Zhengcheng Shen
Max Planck Institute for Plasma Physics, Germany
Harold Soh
Harold Soh
Associate Professor at National University of Singapore
Human Robot InteractionMachine LearningTactile PerceptionArtificial IntelligenceRobotics