Memory-Guided B-Roll Generation from User Video Collections
This study addresses the challenge of extracting visual evidence from massive user videos to generate multi-shot B-roll sequences with consistent characters, scenes, and styles. To this end, it proposes MemComposer, a system employing a three-stage architecture that transforms raw videos into entity-centric structured memories. By integrating natural language instruction planning, conditional retrieval, and iterative generation alongside an automated critique-feedback mechanism, the approach enables precise reference-frame-based generation and consistency verification. This work transcends the limitations of conventional text-to-video paradigms through memory grounding. User preference studies demonstrate that the system achieves 60.0% prompt adherence and up to 92.8% visual alignment, significantly outperforming ungrounded baselines and thereby validating the effectiveness of the proposed memory-grounding strategy.