Complementary Retrieval-Augmented Prompting for Consistent Long-Form Video Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of maintaining cross-shot consistency in long video generation, where existing retrieval methods suffer from information redundancy or insufficient coverage. To this end, this work proposes a Complementary Retrieval-Augmented Prompting framework. By parsing screenplays to construct a visual element registry, the method introduces a novel element-aware complementary retrieval mechanism that intelligently aggregates historical reference images. This provides frozen video generators with comprehensive yet low-noise conditional constraints, enabling training-free coherent long video generation. Experimental results demonstrate that the proposed framework significantly outperforms baseline methods on multi-shot story generation tasks, substantially improving both cross-shot consistency and text controllability while exhibiting high interpretability.
📝 Abstract
While recent video foundation models excel at generating high-quality short videos, long-form video generation remains a critical challenge, where a major bottleneck lies in conditioning independently generated shots to preserve consistent characters, scenes, and objects throughout a story. Existing training-free approaches typically condition target shots using retrieved historical visuals. However, these references often suffer from severe informational mismatch, either introducing irrelevant contextual redundancy or failing to provide the full combination of required elements for the target shot. To resolve this, we present Complementary Retrieval-Augmented Prompting, an agentic framework that strategically aggregates a compact set of mutually supportive historical references to achieve complete and targeted conditioning for long-form video generation without retraining or modifying the underlying generator. Specifically, our framework explicitly models the visual elements required by each target shot by parsing the narrative script into a text-grounded visual element registry that tracks characters, objects, scenes, and their shot-level states. A VLM-annotated keyframe library further maps these elements to past visual observations. Guided by the required elements, our agent retrieves complementary references that maximize target-element coverage while minimizing historical noise. Finally, the retrieved references, structured element states, and grounding instructions are assembled into a unified prompt for the frozen video generator. This element-aware process provides comprehensive conditioning while remaining fully interpretable. Quantitative and qualitative evaluations on multi-shot story generation demonstrate that our method consistently outperforms recent-frame, memory-based, and entity-level retrieval baselines in cross-shot consistency and text-controllability.
Problem

Research questions and friction points this paper is trying to address.

Long-form video generation
Cross-shot consistency
Retrieval-augmented prompting
Informational mismatch
Video foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Prompting
Long-Form Video Generation
Agentic Framework
Visual Element Registry
Cross-shot Consistency
🔎 Similar Papers
No similar papers found.
X
Xianghan Wei
X
Xiaoda Yang
Z
Zhi Wang
A
An Pan
Daoan Zhang
Daoan Zhang
PhD Student, University of Rochester
Computer VisionMultimodal LearningLLM
H
Huayi Zhang
Y
Yan Zhang
W
Wei Xu
Z
Zishun Liao
J
Jianwen Lou