ChordVideo: One-Step, Training-Free, Temporally Consistent Video Editing via Low-Energy Transport

πŸ“… 2026-08-01
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses temporal flickering and editing strength drift in training-free, one-step video editing by extending low-energy smoothing from the sampling-time to the video-time dimension. It introduces motion-aligned causal aggregation and an optional temporally smoothed proximal correction strategy, while establishing a theoretical bound on warping error. Built upon a single-step text-to-image model, the method integrates shared noise injection, low-energy field smoothing, and efficient temporal consistency constraints. Evaluated on the TGVE and DAVIS benchmarks, the approach reduces warping error by 78%, decreases flickering by 49%, improves CLIP-based frame consistency by 9–10 points, and increases background PSNR by approximately 1.5 dBβ€”all with only two network evaluations per frame.
πŸ“ Abstract
One-step text-to-image models enable training-free, inversion-free editing with only 1--2 network function evaluations (NFE), while ChordEdit stabilizes such edits through low-energy smoothing along sampling time. Applied independently to video frames, however, it produces temporal flicker and edit-strength drift. We introduce \textbf{ChordVideo}, which extends the same low-energy principle to video time through shared noise, motion-aligned causal aggregation of per-frame Chord fields, and an optional temporally smoothed proximal correction. We derive a warping-error bound that separates motion bias from stochastic flicker and predicts diminishing returns with larger temporal windows. On TGVE/DAVIS with two one-step backbones, ChordVideo reduces warping error by \textbf{78\%} and flicker by \textbf{49\%}, improves CLIP frame consistency by \textbf{9--10 points}, and increases background PSNR by about \textbf{1.5,dB}, while retaining \textbf{2 NFE/frame}. Compared with seven multi-step editors, it achieves competitive temporal consistency and source preservation using \textbf{10--60$\times$ fewer model steps per clip
Problem

Research questions and friction points this paper is trying to address.

temporal consistency
video editing
temporal flicker
edit-strength drift
one-step editing
Innovation

Methods, ideas, or system contributions that make the work stand out.

ChordVideo
training-free video editing
temporal consistency
low-energy transport
one-step diffusion