Accelerating Video Diffusion via Training-Free Trajectory Routing

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive inference costs of video diffusion models and the bottleneck that step distillation still requires large-model evaluation at every timestep. To overcome these limitations, this work proposes TRACK, a heterogeneous denoising strategy based on trajectory-aware capacity routing and relative discrepancy scoring. By calibrating precise switching points, TRACK dynamically invokes the large model for quality-sensitive steps while routing low-discrepancy steps to a smaller counterpart, establishing a training-free, architecture-agnostic paradigm for dynamic model switching. Experimental results demonstrate that TRACK achieves 1.95× to 2.73× acceleration across multiple mainstream video generation models while preserving comparable visual quality and high output diversity.
📝 Abstract
Video diffusion is computationally expensive, as it requires executing a large model across many denoising steps. Even with step-distillation, inference remains expensive because every distilled step still requires a costly model evaluation. We present TRACK: TRajectory-Aware Capacity routing via top-K selection, a heterogeneous denoising strategy that switches between compatible large and small models at selected steps, reducing the average cost per denoising evaluation. The switching steps are determined using a calibration process. TRACK first rolls out a reference trajectory with the large model. Then at each step, the small model's prediction is also collected and compared against the large model's prediction to obtain a relative disagreement score. Both models receive the same latent, timestep, conditioning, and guidance inputs. Aggregating this signal over a calibration set produces a disagreement score map across diffusion steps, which determines a switching policy for an efficient inference process: quality-sensitive steps keep using the large model, while steps with low disagreement scores are routed to the small model. Inference executes only the selected model at each step, requiring no retraining, architecture or scheduler changes, or online dual-model evaluation. Across Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo, TRACK yields $1.95\times$, $2.04\times$-$2.73\times$, $2.69\times$, and $2.17\times$ speedups, respectively, with comparable aggregate quality and high diversity retention. TRACK thereby establishes automated, training-free model switching as a practical acceleration paradigm for video diffusion.
Problem

Research questions and friction points this paper is trying to address.

Video Diffusion
Computational Cost
Inference Acceleration
Denoising Steps
Innovation

Methods, ideas, or system contributions that make the work stand out.

Training-free acceleration
Trajectory routing
Video diffusion
Heterogeneous denoising
Model switching
🔎 Similar Papers