🤖 AI Summary
This study addresses the intent-foresight inconsistency in world action models arising from the mismatch between perturbed actions and real-world video noise. To this end, we propose an off-path rendering paired training mechanism based on a physics simulator to align perturbed actions with visual signals. Furthermore, we design a video-action collaborative denoising schedule to optimize the noise distribution of diffusion models, and construct a multi-agent interface supporting unified sequence modeling for variable numbers of agents. Experimental results demonstrate that the proposed method significantly enhances action prediction accuracy, video generation consistency, instruction-following capability, and motion fidelity across autonomous driving and robotic tasks.
📝 Abstract
World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/