WAM-Cache: Staleness-Bounded KV Reuse for Efficient World Action Models

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational cost of video DiT prefilling in world action models and the limitation that existing acceleration methods remain dense. We propose a training-free sparse refresh framework that reuses KV caches across blocks, combining action-expert attention with visual surprise to select tokens, recomputing only a sparse refresh set. Furthermore, we reveal that downstream accuracy is dominated by action attention and introduce a strict age bound to suppress error accumulation. Experiments demonstrate that this framework reduces computation by 32–42% in both simulated and real-world robotic tasks, with performance degradation kept within 2.5 percentage points.
📝 Abstract
World Action Models (WAMs) enable generalist robot manipulation by conditioning an action expert on representations from a pretrained video Diffusion Transformer (DiT). In closed-loop control, the video DiT runs at every chunk to encode the current observation into layerwise key-value (KV) pairs that the action expert queries. This prefill dominates the per-chunk computational cost, yet existing training-free accelerations leave it fully dense. We present WAM-Cache, a training-free framework that retains layerwise key-value representations across chunks and recomputes only a sparse refresh set of tokens. Crucially, we find that the intuitive heuristic of refreshing visually drifted tokens plateaus far below the dense baseline, even with an oracle predicting ground-truth KV drift. Downstream action accuracy is instead governed by where the action expert attends, not by what moved. WAM-Cache therefore selects the refresh set by uniting the action expert's cross-attention with visual latent surprise, complemented by a strict age bound that suppresses compounding error. On Fast-WAM, WAM-Cache cuts video DiT prefill FLOPs by 32-42% across RoboTwin 2.0, LIBERO, and real-world experiments, while staying within 0.7-1.8 percentage points of the dense policy in simulation and 2.5 points on a real robot.
Problem

Research questions and friction points this paper is trying to address.

World Action Models
KV cache reuse
computational efficiency
robot manipulation
Diffusion Transformer
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Action Models
KV Cache Reuse
Training-free Acceleration
Cross-attention Guided Refresh
Staleness Bound
🔎 Similar Papers