🤖 AI Summary
This study addresses the inefficient computation in existing world action models that employ fixed denoising steps, failing to accommodate varying sensitivity of actions to generation errors. We propose an adjustable-budget prediction framework that trains an interval-conditioned flow matching model via budget-aligned teacher trajectory distillation and leverages a shared low-rank adapter to support flexible single- to multi-step generation. Additionally, a risk-reward scheduler based on curvature difficulty and fidelity prediction dynamically selects the minimal computational budget satisfying accuracy requirements. Evaluations across three mainstream benchmarks demonstrate that our method reduces denoising steps by 50%–85% on average while maintaining baseline success rates, and improves task success by 7%–12% under single-step inference. The effectiveness is further validated on real-world robotic platforms.
📝 Abstract
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.