Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the fragility of existing video-generation-based World-Action Models (WAMs) under visual distribution shifts, which struggle to balance the benefits of large-scale pretraining with semantic robustness. The authors propose a post-training approach that injects appearance-invariant dynamic semantics into the action prediction stream while preserving the pretrained VAE representation space. This is achieved through a lightweight semantic lookahead alignment mechanism comprising learnable query tokens, positional encoding alignment, and semantic hidden state matching. Notably, the method enables robust control against out-of-domain perturbations—such as lighting changes—without modifying the original generative pathway or requiring re-pretraining. Experiments demonstrate consistent improvements in success rates across multiple WAM baselines on both simulated out-of-distribution generalization benchmarks and real-world robotic tasks, all while maintaining in-distribution performance.
📝 Abstract
Mainstream World-Action Models (WAMs) adapt pretrained video generation models (VGMs) for robot control, transferring their learned dynamics prior for action prediction. These VGMs are typically trained in a variational autoencoder (VAE) latent space. However, the VAE latent space is optimized for pixel reconstruction, which rewards fine appearance detail and leaves the action prediction fragile under visual shifts. Recent works build WAMs in semantic latent space, which are more robust to appearance shifts. However, these models cannot leverage the large-scale VGM pretraining that exists only in VAE space. To overcome this dilemma, we propose Robust-WAM, a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream. This retains the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics that stay reliable under illumination shifts and other visual out-of-distribution conditions. Specifically, we employ learnable query tokens to bring future-scene semantics into the action stream by aligning their output hidden states with the semantic foresight of future ground-truth frames. To establish the temporal correspondence between each query and the future step it describes, we give it the positional encoding of the matching action tokens. Experiments on out-of-distribution generalization simulation benchmarks and a real-robot setup show that our Robust-WAM consistently improves the success rates of multiple WAM baselines without sacrificing in-distribution performance.
Problem

Research questions and friction points this paper is trying to address.

World-Action Models
video generation
visual distribution shift
semantic foresight
out-of-distribution generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

World-Action Models
semantic foresight
VAE latent space
out-of-distribution generalization
post-training alignment
🔎 Similar Papers
No similar papers found.
Haodong Yan
Haodong Yan
PhD student of INTR, HKUST (GZ)
Human reconstructionmotion prediction
J
Junfeng Li
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Junjie He
Junjie He
Guizhou University
MRIDeep LearningCT
Zhide Zhong
Zhide Zhong
Beijing Institute of Technology
Robotics
M
MingMing Yu
Beihang University, Beijing, China
Wenxuan Song
Wenxuan Song
The Hong Kong University of Science and Technology (Guangzhou)
Vision-language-action ModelRobotics
J
Jiaguan Zhu
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Y
Yangyang Zheng
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Y
Yuqiao Du
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
J
Jiadi You
The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China
Y
Yingjie Cai
Huawei Foundation Model Department
X
Xu Yan
Huawei Foundation Model Department
G
Guanyi Zhao
Huawei Foundation Model Department
Bingbing Liu
Bingbing Liu
Researcher, Huawei
Autonomous DrivingRoboticsNeural RenderingVision Foundation Model
Haoang Li
Haoang Li
Assistant Professor, Hong Kong University of Science and Technology (Guangzhou)
Robotics3D Computer Vision