FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis

πŸ“… 2025-05-02
πŸ“ˆ Citations: 1
✨ Influential: 0
πŸ“„ PDF

career value

229K/year
πŸ€– AI Summary
Long video generation suffers from distribution shift induced by frame expansion when adapting short-video-trained diffusion models, resulting in visual inconsistency and motion distortion. To address this, we propose a training-free decoupling framework: first, we identify that PCA precisely separates global appearance consistency from local motion intensity in videos; second, we design a progressive feature fusion strategy and an initial noise mean reuse mechanism to break the strong appearance-motion coupling. Our method is plug-and-play for mainstream diffusion models, requiring only statistical feature reuse and cosine-similarity guidanceβ€”no fine-tuning or parameter updates. Experiments demonstrate significant improvements in long-video visual quality and inter-frame appearance consistency, achieving state-of-the-art performance across multiple benchmarks.

Technology Category

Application Category

πŸ“ Abstract
Long video generation involves generating extended videos using models trained on short videos, suffering from distribution shifts due to varying frame counts. It necessitates the use of local information from the original short frames to enhance visual and motion quality, and global information from the entire long frames to ensure appearance consistency. Existing training-free methods struggle to effectively integrate the benefits of both, as appearance and motion in videos are closely coupled, leading to motion inconsistency and visual quality. In this paper, we reveal that global and local information can be precisely decoupled into consistent appearance and motion intensity information by applying Principal Component Analysis (PCA), allowing for refined complementary integration of global consistency and local quality. With this insight, we propose FreePCA, a training-free long video generation paradigm based on PCA that simultaneously achieves high consistency and quality. Concretely, we decouple consistent appearance and motion intensity features by measuring cosine similarity in the principal component space. Critically, we progressively integrate these features to preserve original quality and ensure smooth transitions, while further enhancing consistency by reusing the mean statistics of the initial noise. Experiments demonstrate that FreePCA can be applied to various video diffusion models without requiring training, leading to substantial improvements. Code is available at https://github.com/JosephTiTan/FreePCA.
Problem

Research questions and friction points this paper is trying to address.

Addressing distribution shifts in long video generation from short videos
Integrating local and global information for visual consistency
Decoupling appearance and motion using Principal Component Analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Decouples appearance and motion via PCA
Integrates global consistency and local quality
Training-free long video generation paradigm
πŸ”Ž Similar Papers
No similar papers found.
J
Jiangtong Tan
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China
H
Hu Yu
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China
J
Jie Huang
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China
Jie Xiao
Jie Xiao
University of Science and Technology of China
low level visiongenerative modelmachine learning
F
Feng Zhao
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China