🤖 AI Summary
This study addresses the high inference latency of video diffusion models and the limitation of existing cross-request reuse methods, which rely on coarse-grained semantic similarity and struggle to balance acceleration with generation quality. We propose a precise reuse method grounded in generative native compatibility. By extracting early probe states to construct signatures, this work directly evaluates computational compatibility from within the model for the first time, guiding risk-aware reuse of historical latents and sparse attention. This effectively resolves the efficiency bottleneck caused by conflating semantic similarity with generative compatibility. The proposed approach achieves up to 2.17× end-to-end speedup across three backbone architectures while maintaining competitive generation quality under strict caching strategies.
📝 Abstract
Video diffusion transformers produce high-quality videos, yet iterative denoising incurs substantial inference latency, limiting interactive and large-scale serving. Most existing acceleration methods focus on individual requests, thereby restricting efficiency gains to redundancy within a single generation trajectory. Recent cross-request reuse offers an additional source of savings, but existing approaches often infer reusability from coarse semantic similarity. This conflates semantic relatedness with generation-level computational compatibility, so aggressive reuse may accept incompatible historical computation while conservative reuse leaves substantial acceleration unrealized. We present \emph{Carnator}, a cross-request acceleration framework that addresses this challenge by extracting and using generation-native compatibility evidence directly from the model's evolving internal states. Specifically, \emph{Carnator} performs a lightweight early probe to construct an Early Signature from internal diffusion states, assessing reuse validity through risk-aware compatibility decisions. The same evidence characterizes reuse scope by localizing target-specific computation and guiding joint reuse of historical latent trajectories and sparse attention connectivity. Across three text-to-video backbones, Carnator consistently achieves higher cache-hit end-to-end acceleration than the evaluated cross-request baselines despite more selective cache acceptance, reaching up to 2.17$\times$ speedup while maintaining competitive generation quality.