CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing caching strategies in interactive video world models, which neglect control signals and consequently incur substantial denoising overhead and sluggish responsiveness. We propose a training-free, control-aware caching framework that introduces a zero-overhead scheduling mechanism based on control sequences. By integrating action-state detection with low-frequency structural persistence analysis, the method optimizes cache reuse. Furthermore, it dynamically adjusts computational scheduling through residual reuse and frequency-mixed historical priors. Evaluated on benchmarks such as Matrix-Game, our approach achieves 1.21× to 1.41× backbone acceleration and significantly improves comprehensive WBench scores, effectively balancing generation efficiency with interaction quality.
📝 Abstract
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising, but each chunk still requires several costly denoising iterations. Training-free caching can reduce this cost, yet existing policies make reuse decisions primarily from model-internal denoising dynamics and do not explicitly account for control transitions. Actually, interactive generation explicitly exposes a signal they do not use: the controls for a chunk arrive before it is denoised, so a schedule derived from them costs no forward pass. To this end, we analyze adjacent chunks under different control regimes and find that structural similarity drops around action changes, while low-frequency structure remains more persistent than high-frequency detail. Motivated by these observations, we propose CtrlCache, a training-free control-aware caching framework that adapts computation to the current control sequence. Specifically, the action-aware scheduling and refresh policy detects action changes across and within chunks, and labels each chunk as initial, transition, turning, or steady state. At one selected interior denoising step, initial and transition chunks retain full computation, while turning and steady chunks reuse the transformer residual from the most recent fully computed step in the same chunk. To exploit the persistence of low-frequency structure during steady interaction, we further introduce a frequency-mixed history prior guidance that incorporates complementary information from the preceding clean latent without an additional DiT forward pass. Evaluated on Matrix-Game 2.0 and LingBot-World v1/v2, CtrlCache achieves 1.21x to 1.41x DiT-backbone speedups without model retraining while improving WBench Overall scores over original inference across all three models.
Problem

Research questions and friction points this paper is trying to address.

Interactive video world models
Autoregressive generation
Denoising acceleration
Training-free caching
Control-aware computation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Control-Aware Caching
Training-Free Acceleration
Action-Aware Scheduling
Frequency-Mixed History Prior
Video World Models