Native Action-Prior Learning from Videos for World Action Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing world action models that rely heavily on annotated trajectories and struggle to scale with unlabeled videos. We propose NAVA-WAM, which introduces a native action prior learning paradigm to pretrain action policies directly from observational videos. Methodologically, we employ a two-stage training procedure to optimize an Action-DiT directly, eliminating dependence on separate latent action models. By integrating flow matching supervision, structured joint attention, and asymmetric attention mechanisms, our approach effectively decouples visual representations from the action denoising process. Experimental results demonstrate that NAVA-WAM outperforms existing baselines on both in-distribution and out-of-distribution tasks, significantly improves label efficiency, and exhibits strong generalization capabilities on real-world robots.
📝 Abstract
World action models integrate future visual dynamics with robot action prediction, but their scalability remains limited by the need for action-annotated robot trajectories. Observation-only videos contain rich evidence about interaction dynamics, but existing approaches typically use them either to pretrain visual representations that must later be adapted for control, or to infer latent actions that are subsequently grounded to robot commands. We present NAVA-WAM, which introduces native action-prior learning by directly pretraining the action policy from observation-only videos, avoiding indirect representation-to-control transfer or a separate latent-action model. Our training consists of two stages. First, we pretrain on observation-only videos, where future-video flow-matching supervision over visual transitions is propagated through transition-structured joint attention to optimize the Action-DiT and learn action-relevant priors. Second, we use action-labeled demonstrations to post-train the Action-DiT for robot control through joint video--action flow matching, while asymmetric attention decouples the visual branch from iterative action denoising and enables efficient action-only inference. Extensive experiments show that NAVA-WAM consistently outperforms prior approaches under both in-distribution and out-of-distribution settings, while demonstrating strong action-label efficiency and effective real-robot generalization. These results establish native action-prior learning as an effective approach to directly pretrain action policies from observation-only videos, providing a scalable path beyond action-labeled robot data.
Problem

Research questions and friction points this paper is trying to address.

world action models
action-prior learning
observation-only videos
scalability
robot control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native Action-Prior Learning
World Action Models
Flow Matching
Action-DiT
Asymmetric Attention
🔎 Similar Papers
No similar papers found.