🤖 AI Summary
This study addresses the inefficiency and poor reusability of existing video diffusion model distillation, which requires repetitive training for each downstream task. We propose a modular "train once, reuse everywhere" distillation framework that encapsulates capabilities such as accelerated sampling, classifier-free guidance (CFG), and long-video error correction into LoRA adapters. These adapters enable plug-and-play deployment, facilitating cross-architecture and cross-conditioning transfer without retraining, while supporting training-free inference weight modulation and flexible composition of multiple capabilities. Extensive experiments across 54 downstream models and eight task categories demonstrate that our approach substantially reduces training costs while preserving superior generation quality and controllability.
📝 Abstract
Video diffusion models are increasingly developed into specialized models for diverse downstream tasks, and this development often includes a distillation stage, for example to accelerate sampling or to improve long-video generation. This stage is typically repeated for every specialized model. We introduce LongLive-Plug, a once-for-all distillation framework that learns reusable capabilities as LoRAs on a base model for training-free, plug-and-play deployment to compatible downstream models. These capabilities include single-pass classifier-free guidance, few-step sampling, and long-context error correction for autoregressive generation. The adapters remain reusable even when downstream models add conditioning branches, expand output channels. Despite training at a fixed guidance scale, our dedicated CFG LoRA provides text guidance control through its inference weight. Combining it with a few-step LoRA simultaneously preserves few-step generation and CFG controllability on downstream tasks. We verify training-free deployment on 54 downstream models across three backbone families and eight task categories, including world modeling, robotics, editing, and multimodal generation. The approach may support additional compatible models. Each capability can thus be distilled once per backbone family and reused without per-target retraining.