From Scores to Samples: Elastic Forcing for Autoregressive Video Generation

📅 2026-09-28
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the computational burden and architectural constraints imposed by reliance on bidirectional diffusion teachers and online score models in few-step autoregressive video generation. To this end, it proposes a direct distribution matching method that eliminates auxiliary model dependencies during post-training by learning distributions directly from reference videos via maximum mean discrepancy minimization. The approach integrates a frozen self-supervised representation space, a hybrid Nystrom–Monte Carlo estimator, memory-efficient replay, and gradient subsampling. Experiments demonstrate that this framework enables the acquisition of novel visual styles and semantic concepts from reference videos. Notably, a 1.3B-parameter model achieves an improved VBench score of 84.64 while maintaining 17 FPS, and efficient post-training of a 14B-parameter model is realized on eight H200 GPUs.
📝 Abstract
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework minimizes maximum mean discrepancy (MMD) in frozen self-supervised video representation spaces, using a hybrid Nystr\"om--Monte Carlo estimator to balance approximation bias and sampling variance. Memory-efficient replay and gradient subsampling make this objective practical. Using the same architecture and initialization as Self-Forcing, our 1.3B model improves the VBench Total score from 83.80 to 84.64 while retaining 17 FPS. Removing auxiliary score models also enables 14B post-training on eight H200 GPUs. Beyond distillation, learning from reference videos enables the acquisition of new visual styles, semantic concepts, and spatial priors without a target-specific diffusion teacher.
Problem

Research questions and friction points this paper is trying to address.

autoregressive video generation
distribution matching distillation
score models
few-step generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Autoregressive Video Generation
Maximum Mean Discrepancy
Distribution Matching Distillation
Self-supervised Representation
Few-step Generation