TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting

📅 2026-07-30
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a novel paradigm for text-driven video re-rendering that eliminates the need for explicit 3D priors or multi-view paired data, addressing the limited generalization and inability of existing methods to synthesize unseen regions. By leveraging semantic viewpoint specification and timestep-aware data augmentation, the approach enables flexible control over shot scale, camera angle, and narrative perspective. A key insight is that camera motion is predominantly established during high-noise diffusion timesteps, which informs a self-supervised training strategy that obviates reliance on paired data or 3D reconstruction. Integrating timestep-sensitive analysis with joint text-camera conditioning, the method significantly outperforms prior art in camera control accuracy and temporal consistency, supporting large-scale viewpoint transitions, reverse re-rendering, and plausible extrapolation beyond the source view.
📝 Abstract
Video re-shooting aims to regenerate videos with controllable camera motion and viewpoint. Existing methods rely on explicit 3D priors, which are limited by reconstruction quality and often perform poorly when synthesizing previously unseen regions, or on paired videos with different camera trajectories, whose scarcity hinders generalization. We revisit video re-shooting through text-driven semantic viewpoint specification, enabling control over shot scale, viewing angle, and first-/third-person perspective. To this end, we propose TARS, a 3D-free video re-shooting paradigm. Timestep-wise sensitivity analysis reveals that camera motion is primarily established during high-noise stages, where coarse spatiotemporal structures are formed. Based on this insight, we introduce self-supervised training to learn camera dynamics and fundamental visual representations without paired re-shooting data or 3D reconstruction. Through data scaling and joint textual-camera conditioning, TARS supports robust camera and viewpoint control, plausibly synthesizing regions beyond the source view under large camera motions while enabling reverse-angle re-shooting and perspective switching. Extensive experiments show that TARS provides more accurate and temporally consistent camera control than prior methods. Project Page: https://ymlinfeng.github.io/TARS.github.io/
Problem

Research questions and friction points this paper is trying to address.

video re-shooting
3D-free
camera motion
viewpoint control
unseen region synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

3D-free video re-shooting
timestep-aware scaling
text-driven viewpoint control
self-supervised camera dynamics
joint textual-camera conditioning