🤖 AI Summary
This study addresses the high iterative inference latency and state-reuse fragility inherent in world action models by proposing WAMACHINE, a training-free acceleration framework. WAMACHINE introduces a pioneering three-level state adaptation mechanism spanning planning horizons, denoising steps, and Transformer layers, achieving adaptive adjustment of intermediate states through trajectory remapping, observation rebinding, and residual scaling. Evaluated on benchmarks such as LIBERO, the proposed method delivers 1.47–3.05× end-to-end speedup while preserving over 96% of the original task success rate. By effectively reconciling inference efficiency with control precision, this work offers a practical solution for accelerating diffusion- and flow-based world action models without compromising their performance in downstream robotic tasks.
📝 Abstract
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.