AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficient computation in existing world action models that employ fixed denoising steps, failing to accommodate varying sensitivity of actions to generation errors. We propose an adjustable-budget prediction framework that trains an interval-conditioned flow matching model via budget-aligned teacher trajectory distillation and leverages a shared low-rank adapter to support flexible single- to multi-step generation. Additionally, a risk-reward scheduler based on curvature difficulty and fidelity prediction dynamically selects the minimal computational budget satisfying accuracy requirements. Evaluations across three mainstream benchmarks demonstrate that our method reduces denoising steps by 50%–85% on average while maintaining baseline success rates, and improves task success by 7%–12% under single-step inference. The effectiveness is further validated on real-world robotic platforms.
📝 Abstract
World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
Adaptive Inference
Computational Budget
Denoising Steps
Robotic Manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
Budget-Aligned Distillation
Adaptive Inference
Low-Rank Adapters
Risk-Benefit Scheduler
🔎 Similar Papers
No similar papers found.