Efficient World Action Model Inference with Adaptive Intermediate States

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high iterative inference latency and state-reuse fragility inherent in world action models by proposing WAMACHINE, a training-free acceleration framework. WAMACHINE introduces a pioneering three-level state adaptation mechanism spanning planning horizons, denoising steps, and Transformer layers, achieving adaptive adjustment of intermediate states through trajectory remapping, observation rebinding, and residual scaling. Evaluated on benchmarks such as LIBERO, the proposed method delivers 1.47–3.05× end-to-end speedup while preserving over 96% of the original task success rate. By effectively reconciling inference efficiency with control precision, this work offers a practical solution for accelerating diffusion- and flow-based world action models without compromising their performance in downstream robotic tasks.
📝 Abstract
World Action Models (WAMs) enable future-aware control by jointly modeling actions and environment dynamics. However, iterative diffusion or flow inference incurs substantial denoising latency. Prior inference state offers a natural opportunity for acceleration, yet changing planning contexts, observations, and intermediate representations can quickly render retained state stale. Preserving useful computation therefore requires adapting inference state rather than reusing it as-is. To this end, we present $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$, a training-free framework that accelerates WAM inference by preserving and adapting inference state for efficient and accurate continuation as the control loop evolves. Across closed-loop replans, Trajectory Remapping remaps replan state from the preceding replan to initialize the next replan, reducing redundant trajectory generation. Across denoising steps, Observation Rebinding performs anticipatory inference during action execution and rebinds retained denoising state to the real observation for continuation when consistency checks pass, reducing latency exposed to the control loop. Across Transformer layers, Residual Rescaling selectively rescales retained layer state and refreshes it through full computation of the middle layers when probe checks fail, reducing repeated Transformer computation. Evaluations of three representative WAM architectures on LIBERO and RoboTwin 2.0 show that $\mathrm{WAM}{\scriptstyle\mathrm{ACHINE}}$ achieves 1.47-3.05$\times$ speedups in observation-to-action latency and 2.23-3.27$\times$ speedups in GPU inference time per replan, while preserving 96.69-99.54% of native WAM task success.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Inference Acceleration
Denoising Latency
Adaptive Intermediate States
Control Loop
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Training-free Inference Acceleration
Trajectory Remapping
Observation Rebinding
Residual Rescaling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhinnan Liu
Xiamen University
H
Haozhi Han
Institute for AI Industry Research, Tsinghua University
R
Ruge Zhang
Institute of Computing Technology, Chinese Academy of Sciences
T
Teng Ma
Alibaba Group
T
Tao Ma
Alibaba Group
Z
Zheng Liu
Alibaba Group
Y
Yifeng Chen
School of Computer Science, Peking University
Yunquan Zhang
Yunquan Zhang
Professor of Institute of Computing Technology, CAS
parallel computingparallel programmingparallel computational model
T
Ting Cao
Institute for AI Industry Research, Tsinghua University
Yunxin Liu
Yunxin Liu
IEEE Fellow, Guoqiang Professor, Institute for AI Industry Research (AIR), Tsinghua University
Mobile ComputingEdge ComputingAIoTSystemNetworking
K
Kun Li
Institute for AI Industry Research, Tsinghua University