PhysWAM: Physically Consistent World Action Model for Autonomous Driving

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of shared geometric constraints between scene prediction and action planning in world action models by proposing a unified generative framework. The core innovation is a coupled point projection mechanism that directly leverages geometric relationships to align generated depth, ego-vehicle motion, and LiDAR through label-free consensus rules, replacing complex scorers to achieve physically consistent joint generation. Built upon a flow matching Transformer architecture, the model supports joint denoising of multi-view videos, metric depth, and SE(3) motions. Experimental results demonstrate that the proposed method achieves superior planning performance on the NAVSIM benchmark, enables zero-shot closed-loop transfer, and significantly improves both depth prediction accuracy and video temporal coherence.
📝 Abstract
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
Problem

Research questions and friction points this paper is trying to address.

Autonomous Driving
World Action Model
Physical Consistency
Geometric Constraint
Metric Depth
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Model
Flow-Matching Transformer
Coupled Point Projection
Autonomous Driving
Physical Consistency
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.