🤖 AI Summary
This study addresses the challenge of extracting visual evidence from massive user videos to generate multi-shot B-roll sequences with consistent characters, scenes, and styles. To this end, it proposes MemComposer, a system employing a three-stage architecture that transforms raw videos into entity-centric structured memories. By integrating natural language instruction planning, conditional retrieval, and iterative generation alongside an automated critique-feedback mechanism, the approach enables precise reference-frame-based generation and consistency verification. This work transcends the limitations of conventional text-to-video paradigms through memory grounding. User preference studies demonstrate that the system achieves 60.0% prompt adherence and up to 92.8% visual alignment, significantly outperforming ungrounded baselines and thereby validating the effectiveness of the proposed memory-grounding strategy.
📝 Abstract
We introduce an approach for collection-grounded B-roll sequence generation. Given a user's video collection, a directive given in natural language, and a target duration, the goal is to produce a multi-shot sequence that complements the user's primary footage (A-roll) while preserving the collection's characters, settings, objects, and style. This task is challenging as one must choose the visual evidence from hours of captured footage that should guide the generation of each shot in the sequence. We address this challenge with MemComposer, a three-stage system that turns raw footage into a structured memory with visual references (characters, settings, objects, and style) and uses it to plan, retrieve, and generate grounded B-roll sequences. First, in a one-time offline stage, MemComposer constructs an entity-centric memory from raw video. Second, it uses the memory and user directive to plan a grounded sequence and retrieve conditioning frames for each shot. Third, it iteratively generates and critiques the sequence to enforce identity, setting, and sequence-level consistency. We evaluate MemComposer in a user preference study along two dimensions: prompt adherence and visual alignment to the user's collection. Against an ungrounded text-to-video planner, MemComposer wins 60.0\% of prompt-adherence and 92.8\% of visual-alignment comparisons, showing the grounding benefit of collection memory and reference retrieval. Against retrieval-only sequences assembled from captured footage, MemComposer wins 94.5\% of prompt-adherence comparisons, showing the value of generating missing shots, while retrieval-only sequences are preferred for visual alignment in 58.2\% of comparisons.