🤖 AI Summary
This study addresses the issue of irrelevant content intrusion when composing multiple historical references for new shots in interactive long video generation. To this end, a training-free framework is proposed. Methodologically, it introduces a pioneering LLM-based semantic slot routing mechanism to enable component-level precise retrieval, alongside a masked memory weaving technique that isolates extraneous visual information. Furthermore, an overlap-adaptive RoPE is designed to dynamically adjust temporal offsets, thereby eliminating artifacts caused by missing references. This work significantly enhances cross-shot subject and background consistency. By maintaining high visual quality and text alignment, it effectively resolves the generation degradation induced by the blending of historical references.
📝 Abstract
Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.