DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the memory and computational bottlenecks caused by growing KV caches in autoregressive video diffusion models. We propose a training-free dynamic cache compression method that formulates cache pruning as a denoising consistency problem, leveraging intrinsic consistency signals to guide token selection and preserve highly divergent tokens for maintaining long-range temporal information. Additionally, we introduce CMBench, a dedicated evaluation benchmark for this task. Experimental results demonstrate that our approach achieves an 85.43% cache reduction and a 4.14× generation speedup while maintaining DINO scores comparable to full-cache baselines. By effectively balancing inference efficiency with generation quality without requiring additional training, this work offers a practical solution for scalable video generation.
📝 Abstract
Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compression strategies discard history using fixed windows or select tokens through local attention and similarity signals, without directly measuring whether a chunk contributes information beyond the retained context. We introduce DeCoPrune, a training-free method that treats cache compression as a denoising-consistency problem. We find that tokens with larger discrepancies between intermediate clean predictions and final denoised values tend to carry visual evidence less predictable from the retained context. DeCoPrune uses this model-intrinsic signal to retain high-discrepancy tokens in the long-term cache while pruning low-discrepancy tokens. To evaluate information retention, we introduce CMBench, comprising 58 approximately one-minute generated or real-world context episodes and 116 Reappear or Revisit continuation tasks requiring recall of earlier events or objects. Experiments with LingBot World v2 show that DeCoPrune achieves a DINO score of 0.6701 on a 0-1 scale, with an 85.43% reduction in cumulative historical KV token counts and a 4.14-fold continuation-generation speedup over FullKV. Its head-specialized variant reaches 0.6783 at an 86.19% pruning ratio, approaching FullKV's 0.6803 score and exceeding the evaluated compression baselines at similar budgets. These results indicate that denoising consistency can support long-range information retention while reducing autoregressive inference cost. Our project homepage is https://decoprune.github.io. The code is available at https://github.com/DeCoPrune/CMBench, and the benchmark at https://huggingface.co/datasets/Aoraku/CMBench.
Problem

Research questions and friction points this paper is trying to address.

Autoregressive video diffusion
KV cache pruning
Cache compression
Long-range information retention
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-Cache Pruning
Denoising Consistency
Autoregressive Video Diffusion
Training-free Compression
CMBench