Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the safety challenge of predicting action executability in robotic social navigation by proposing a latent world model based on first-person videos. By formulating an affordance-aware inverse dynamics objective, the method directly learns the discrepancy between nominal and actual actions as a safety signal, enabling secure planning without privileged information or online reinforcement learning. The proposed approach supports zero-shot transfer, achieving a 63.77% success rate on the Social-HM3D benchmark while reducing human collision rates to 21.67%. Furthermore, it maintains strong performance on Social-MP3D. Overall, this work presents an efficient and safe paradigm for vision-driven social navigation.
📝 Abstract
Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

social navigation
robot safety
latent world model
action realizability
collision avoidance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent World Model
Social Navigation
Inverse Dynamics
Safety-Aware Planning
Realizable Action
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhihao Zheng
Computer Science and Engineering department, P.C. Rossin College of Engineering and Applied Science, Lehigh University, Bethlehem, PA 18015, USA
Mooi Choo Chuah
Mooi Choo Chuah
Professor of CSE Dept, Lehigh University
mobile computingmobile healthcarecomputer visionnetwork securitydisruption tolerant