🤖 AI Summary
This study addresses the degraded generalization of implicit world action models caused by discarding future predictions during inference. We reveal that the initial denoising step is critical for generalization, as performance is determined by future representation preparation rather than the generation process itself. Accordingly, we propose Simple-WAM, which simplifies future modeling to a single fully-noised forward pass with an adjusted noise schedule, thereby balancing the strengths of explicit and implicit models. Built upon a video diffusion architecture, our method employs backbone-matched training and test-time future modeling techniques. Experiments demonstrate that Simple-WAM achieves generalization performance surpassing explicit models in both simulated and real-world tasks, while maintaining inference efficiency comparable to implicit models.
📝 Abstract
World action models (WAMs) predict the future alongside actions during \emph{training}. Due to the heavy computation cost of video denoising, whether the future must still be generated during \emph{inference} is disputed: Explicit WAMs denoise it into clean frames along with every action chunk, whereas Latent WAMs discard it entirely for acceleration. We find that latent WAMs, despite matching explicit ones on in-distribution tasks, fail to retain the generalization benefits that originally motivated WAMs. To demonstrate this, we evaluate generalization along three axes: \emph{environmental perturbation}, \emph{data efficiency}, and \emph{task generalization}. Controlled comparisons with a matched backbone, training data, and budget reveal consistent degradation across all three axes when the action expert no longer conditions on future representations. Further analysis shows that the gap arises almost entirely from the first denoising step: the benefit comes from \emph{preparing} the future, not \emph{generating} it. We therefore propose \textbf{Simple-WAM}, which simplifies future modeling into a single forward pass of fully noised video tokens and adapts the training-time noise schedule to this inference behavior. Across simulation and real-world tasks, Simple-WAM achieves the best of both worlds, leading explicit WAMs in generalization performance with efficiency comparable to Latent WAMs. Project Page: \href{https://zrporz.github.io/Simple-WAM-Web/}{\textcolor{panton}{\texttt{https://zrporz.github.io/Simple-WAM-Web}}}