Connected Self Forcing: Beyond Local Learning in Video Autoregression

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of local learning and long-term temporal inconsistency in autoregressive video generation caused by gradient truncation. To overcome these issues, we propose a Connected Self-Forcing training framework that reconnects cross-chunk gradient pathways, enabling feedback from subsequent predictions to guide earlier context generation and thereby facilitating global dependency modeling. Furthermore, this framework incorporates shortcut gradient replay, distribution matching distillation, and KV caching techniques to optimize streaming generation. The proposed method significantly enhances both the visual quality and temporal consistency of long videos while requiring no modifications to the inference pipeline.
📝 Abstract
To stream long videos while maintaining visual quality and temporal consistency, Self Forcing mitigates exposure bias through self-rollout training on self-generated histories with key-value (KV) caching. To keep memory manageable, it detaches historical caches, preserving forward dependencies between chunks but severing the backward gradient paths. We introduce Connected Self Forcing, a training framework that reconnects gradient paths across autoregressive chunks, allowing feedback from later predictions to guide how earlier context is generated. These connections go beyond historical KV-writing: gradients pass through generated latents into the computations that produced them, linking the generation of earlier context to its use in later predictions. To make this connected training memory-efficient, we develop shortcut gradient replay, which recovers cross-chunk gradients without retaining the full rollout computation graph. Integrated with distribution matching distillation, Connected Self Forcing trains historical chunks according to both their direct supervision and their contribution to subsequent generation. Experiments on autoregressive video generation show improvements in long-horizon visual quality and temporal consistency, without changing the inference procedure.
Problem

Research questions and friction points this paper is trying to address.

autoregressive video generation
exposure bias
temporal consistency
gradient disconnection
long-horizon video
Innovation

Methods, ideas, or system contributions that make the work stand out.

Connected Self Forcing
Video Autoregression
Shortcut Gradient Replay
Distribution Matching Distillation
Exposure Bias
🔎 Similar Papers
2024-01-15IEEE Transactions on Information Forensics and SecurityCitations: 0
Dongbin Zhang
Dongbin Zhang
Tsinghua University
C
Chaoda Zheng
XPeng
Kangjie Chen
Kangjie Chen
Nanyang Technological University
Trustworthy AIRed-teamingBackdoor AttacksLLM-based Agents
X
Xiangyu Li
XPeng
S
Shijia Chen
XPeng
J
Jinhao Deng
XPeng
Y
Yuqi Zhang
XPeng
G
Guangfeng Jiang
XPeng
H
Hongbin Lin
The Chinese University of Hong Kong
C
Choo Sin Wai
Tsinghua University
M
Minqi Wang
The Chinese University of Hong Kong
Puyi Wang
Puyi Wang
CUHK CSE PhD
J
Jingye Zhang
Tsinghua University
Y
Yu Zhang
XPeng
X
Xianming Liu
XPeng
B
Boyang Wang
XPeng