MORCA: Offline-to-Online Reinforcement Learning for Adaptive Cache Reuse in Video Diffusion Acceleration

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high inference latency of video diffusion models and the limitations of existing caching methods, which suffer from inaccurate error estimation and imprecise acceleration control. To this end, we propose MORCA, a framework that enables adaptive cache reuse via offline-to-online reinforcement learning. By establishing the first connection between latent variables and terminal errors, MORCA overcomes the inherent constraints of step-wise error estimation. Furthermore, it introduces a latent-aware decision mechanism and a dynamic scheduling algorithm to achieve precise control under user-specified acceleration ratios. Extensive experiments demonstrate that, under equivalent computational budgets, MORCA significantly improves generation fidelity across multiple video generation models, outperforming state-of-the-art methods.
📝 Abstract
Diffusion Transformers (DiTs) achieve remarkable performance in video synthesis, but their iterative denoising process suffers from high inference latency. To address this, caching has emerged as an effective acceleration strategy by capitalizing on inter-step redundancy during denoising. Existing dynamic caching methods typically estimate the error that cache reuse would introduce at each denoising step (step error) to guide cache decisions, whereas our concern is how much quality loss cache reuse would cause in the final generated video (terminal error). We show that step error does not directly correspond to terminal error and that latent information helps capture their relationship, thereby informing cache decisions. Moreover, existing threshold-based methods cannot provide precise speedup control, making it difficult to meet practical requirements for user-specified acceleration targets. To address these limitations, we introduce MORCA, a cache scheduling framework trained through offline-to-online reinforcement learning to make latent-aware reuse/recompute decisions under user-specified acceleration targets. Extensive experiments on different video generation models across multiple target acceleration ratios demonstrate that MORCA achieves better generation fidelity than state-of-the-art caching methods under comparable computational budgets. Code is available at https://github.com/x10ngyx/MORCA.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion Acceleration
Cache Reuse
Diffusion Transformers
Inference Latency
Terminal Error
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline-to-Online Reinforcement Learning
Adaptive Cache Reuse
Video Diffusion Acceleration
Terminal Error
Diffusion Transformers