🤖 AI Summary
This study addresses the detail degradation in few-step distillation of video diffusion models, where student models struggle to match the high-curvature trajectories of teachers. To overcome this, we propose parameterized trajectory distillation, which models teacher trajectory segments as polynomials and guides students to learn teacher supervision along their own predicted paths, enabling curvature to adaptively align with backbone capacity. Combined with LoRA fine-tuning, the method incurs no additional modules during inference. Our approach achieves state-of-the-art four-step generation on Wan2.1 and MiniMax-H3, significantly outperforming PDD and LightX2V Turbo while effectively preserving motion diversity. Human preference evaluations demonstrate win rates of 55.1% and 63.4%, reflecting substantial improvements in dynamic quality and naturalness.
📝 Abstract
Video diffusion and flow models require many sequential evaluations, making generation computationally expensive. Few-step distillation reduces this cost but poses a capacity allocation problem: a student must match the teacher's iterative generation with far less sequential computation. Existing trajectory methods ask the student to reproduce teacher transitions that are highly curved at high noise, which can exceed its capacity and degrade fine detail. We introduce Parametric Trajectory Distillation (PTD), which lets the student parameterize teacher trajectory segments as polynomials and learn from teacher guidance along its own predicted path. PTD is designed to let the learned curvature adapt to the backbone's predictive capacity, preserving motion and diversity. The curvature head is used only in training; inference keeps the original backbone architecture. On Wan2.1-14B, four-step PTD sets a new state of the art for trajectory distillation, significantly improving dynamic quality and naturalness over PDD, the best-performing trajectory-only method on this model, under the same training setting. On the 33B audio-video MiniMax-H3, LoRA-trained PTD significantly improves diversity and naturalness over the state-of-the-art LightX2V Turbo. Blinded human votes give PTD 55.1% and 63.4% preference shares against PDD and LightX2V Turbo. Project page: https://alan-lanfeng.github.io/PTD/.