CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

📅 2026-07-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing text-to-video diffusion models struggle to simultaneously achieve multi-shot structure, fine-grained controllability, and long-term consistency when generating cinematic-length videos. This work proposes a unified framework that requires no retraining: by manipulating positional encodings and attention patterns, it overcomes the temporal continuity bias inherent in pretrained models to enable crisp shot transitions. Furthermore, the method introduces a shot-routing reference conditioning mechanism for per-shot semantic control and an anchor memory mechanism to preserve global visual consistency. For the first time, this approach enables high-quality, long-duration, multi-shot video generation with reference-based controllability—all without model fine-tuning—significantly outperforming existing methods that rely on customization or retraining.
📝 Abstract
Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation across extended temporal horizons. Existing methods rely on customization and retraining to separately address specific requirements, and cannot simultaneously fulfill all the requirements with a unified framework. In this paper, we shed light on the training-free paradigm with the key insight that the difficulty of multi-shot generation arises from a structural bias toward temporal continuity in pretrained video diffusion models, and consequently, propose a unified framework named CineWeaver to achieve reference-controllable multi-shot long-video generation without retraining. We manipulate positional encoding and attention patterns to break temporal continuity during inference to enable clear shot transitions using pretrained video diffusion models. Furthermore, we extend the proposed framework with a shot-routed reference conditioning mechanism for per-shot fine-grained controllability, and develop an anchor memory mechanism to allow long-form generation with consistent global appearance cues. To our best knowledge, CineWeaver is the first unified framework to simultaneously enable \textbf{long-form}, \textbf{reference-controllable}, and \textbf{multi-shot} video generation in a training-free fashion. Experimental results demonstrate that CineWeaver produces high-quality cinematic videos of long durations with consistent identities, stable global appearance, and clear shot transitions. The project page is available at: https://cineweaver.github.io.
Problem

Research questions and friction points this paper is trying to address.

multi-shot video generation
reference-controllable generation
long-form video generation
cinematic storytelling
temporal continuity
Innovation

Methods, ideas, or system contributions that make the work stand out.

training-free
multi-shot video generation
reference-controllable
long-form video
temporal continuity breaking
🔎 Similar Papers
No similar papers found.