Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos

📅 2026-05-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the degradation in generation quality in training-free long video synthesis caused by the mismatch between training and inference paradigms and insufficient long-term temporal coherence. To this end, the authors propose MIGA, a frame-level autoregressive framework that introduces a novel two-stage alignment mechanism to bridge the training-inference gap. MIGA further incorporates a dual consistency enhancement strategy comprising self-reflective correction and long-range low-noise frame guidance. Remarkably, without any additional training, MIGA achieves state-of-the-art performance on the VBench and NarrLV benchmarks and enables high-consistency video generation of unlimited length under constant memory constraints.
📝 Abstract
Without incurring significant computational overhead, train-free long video generation aims to enable foundation video generation models to produce longer videos. Frame-level autoregressive frameworks, e.g., FIFO-diffusion, offer the advantage of generating infinitely long videos with constant memory consumption. However, the mismatch between training and inference, coupled with the challenge of maintaining long-term consistency, limits the effective utilization of foundation models. To mitigate these concerns, we propose \textbf{MIGA}, a novel infinite-frame long video generation method. Firstly, we propose an effective two-stage alignment mechanism that mitigates the training-inference gap by reducing the excessive noise span fed to the model. We then introduce an innovative dual consistency enhancement mechanism, where the self-reflection approach corrects early high-noise frames and the long-range frame guidance approach leverages later low-noise frames with broad coverage to steer generation, jointly improving temporal consistency. Extensive experiments on VBench and NarrLV demonstrate the state-of-the-art performance of MIGA. Our project page is available at https://xiaokunfeng.github.io/miga_homepage/.
Problem

Research questions and friction points this paper is trying to address.

train-free
infinite-frame generation
long video generation
temporal consistency
training-inference gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

train-free video generation
infinite-frame generation
temporal consistency
two-stage alignment
dual consistency enhancement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Xiaokun Feng
Xiaokun Feng
Institute of Automation,Chinese Academy of Sciences
computer versiondeep learning
J
Jiashu Zhu
AMAP, Alibaba Group, Beijing, China
Meiqi Wu
Meiqi Wu
the University of Chinese Academy of Sciences
Computer vision
Chubin Chen
Chubin Chen
Tsinghua University
Generative AI
Fangyuan Mao
Fangyuan Mao
Student at Institute of Computing Technology
Haiyang Guo
Haiyang Guo
Institute of Automation, Chinese Academy of Sciences
Continual LearningMultimodal LearningPattern Recognition
Jiahong Wu
Jiahong Wu
Alibaba-AMAP
AIMLAIGCMLLM
X
Xiangxiang Chu
AMAP, Alibaba Group, Beijing, China
K
Kaiqi Huang
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China; The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China