Edit-VAR: Taming Visual Autoregressive Model for Precise Video Editing

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Edit-VAR,一种无需训练和反演的文本引导视频编辑框架,通过预训练视觉自回归模型实现精准编辑,同时保持源内容和时间一致性。
📝 Abstract
Text-guided video editing modifies target content while preserving the appearance and temporal coherence of unedited regions. Training-based approaches provide strong control but demand substantial data and computation. Training-free methods fall into inversion-free and inversion-based paradigms. Inversion-free approaches avoid trajectory recovery, but their source-preserving guidance can limit editing strength and leave semantic changes incomplete. Inversion-based approaches recover a latent trajectory before regeneration, where approximation errors can accumulate and cause source-content drift and temporal inconsistency. We introduce Edit-VAR, the first training-free and inversion-free framework for text-guided video editing with a pretrained visual autoregressive video model. Edit-VAR directly encodes the source video into multi-scale discrete tokens and performs probability-guided conditional token replacement for source preservation. Attention-guided token-wise and scale-aware modulation selectively relaxes source constraints over edit-relevant positions and generation stages. Scale-Decoupled Generation, implemented as late-scale constraint release, regenerates motion-consistent details and reduces texture fragmentation. Residual-guided token pruning further exploits redundancy at the final two high-resolution scales to reduce inference cost. Extensive experiments and a blind user study demonstrate that Edit-VAR outperforms existing training-free video editing methods overall in editing fidelity, source preservation, temporal coherence, and inference efficiency.
Problem

Research questions and friction points this paper is trying to address.

text-guided video editing
temporal coherence
inversion-free
inversion-based
source preservation
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
inversion-free
probability-guided conditional token replacement
scale-aware modulation
scale-decoupled generation
🔎 Similar Papers
No similar papers found.
C
Chongbo Zhao
Sun Yat-sen University
J
Jiangming Wang
Sun Yat-sen University
X
Xilai Wang
South China University of Technology
X
Xinyu Wang
Tsinghua University
J
Jingyi Tang
Shandong University
C
Chunjie Hao
Nankai University
P
Pengjie Song
Hunan University
Yue Ma
Yue Ma
Bytedance
NLPDialogue SystemLLM