Uncertainty-Aware Consistency Distillation for Few-Step Video Generation

πŸ“… 2026-09-30
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the high latency of multi-step video generation and the unreliability of teacher supervision signals in dynamic regions during conventional distillation by proposing an uncertainty-aware consistency distillation method. The authors reveal that supervision reliability depends on local temporal variation rather than semantic complexity, and accordingly introduce a parameter-free estimator to adaptively relax penalties in high-uncertainty regions. Furthermore, the approach integrates a dual-perturbation consensus objective, exponential weight modulation, feature-space adversarial training, and LoRA fine-tuning to enable efficient few-step generation. Experimental results demonstrate that the proposed method achieves state-of-the-art performance on VBench 2.0 with only four inference steps, attaining a score of 0.556. User studies further confirm its significant superiority over existing approaches.
πŸ“ Abstract
We study few-step video generation, i.e., distilling a multi-step video generator, which typically requires tens of sampling steps, incurring substantial latency and compute, into a few-step student. Consistency distillation is a common recipe, in which a multi-step teacher provides the consistency targets for a few-step student. However, these teacher-guided targets are not equally trustworthy, and the content is harder to learn where it varies rapidly over time, e.g., moving foliage shadows or flowing water. We observe that supervision reliability follows the local difficulty of the content rather than semantic complexity: regions that change little yield consistent endpoint predictions, whereas regions with large temporal variation produce larger discrepancies that coincide with the largest perceptual errors. Motivated by this observation, we propose Uncertainty-Aware Consistency Distillation (UACD), which reweights consistency supervision at each spatiotemporal region using a local, parameter-free uncertainty estimate. Specifically, we construct two independently perturbed teacher-guided consistency paths, whose student endpoint predictions provide a consensus target; the discrepancy between the student's direct prediction and this target is the uncertainty proxy. We then relax the consistency penalty on high-uncertainty regions through an exponential weight, while keeping the full penalty elsewhere, since the student cannot be expected to match targets that are hard to learn. To preserve perceptual quality under aggressive step reduction, we integrate feature-space adversarial training with semantic alignment. With parameter-efficient LoRA adaptation of the 50-step Wan model, our method achieves state-of-the-art 4-step generation on VBench 2.0 (0.556 mean score) and is preferred over competing methods in a user study.
Problem

Research questions and friction points this paper is trying to address.

few-step video generation
consistency distillation
uncertainty-aware
knowledge distillation
video generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-Aware Consistency Distillation
Few-Step Video Generation
Consistency Distillation
LoRA Adaptation
Adversarial Training
πŸ”Ž Similar Papers