Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue of irrelevant content intrusion when composing multiple historical references for new shots in interactive long video generation. To this end, a training-free framework is proposed. Methodologically, it introduces a pioneering LLM-based semantic slot routing mechanism to enable component-level precise retrieval, alongside a masked memory weaving technique that isolates extraneous visual information. Furthermore, an overlap-adaptive RoPE is designed to dynamically adjust temporal offsets, thereby eliminating artifacts caused by missing references. This work significantly enhances cross-shot subject and background consistency. By maintaining high visual quality and text alignment, it effectively resolves the generation degradation induced by the blending of historical references.
📝 Abstract
Recent advances in autoregressive video generation have improved temporal consistency over extended durations, yet interactive storytelling requires more than continuous scene extension: a new shot may combine characters and backgrounds from different historical shots. Whole prompt retrieval can overlook the distinct reference needs of individual components, while directly combining all historical memories may introduce unrelated visual content. To address these problems, we present Weave Forcing, a training-free framework for compositional memory reuse in interactive long video generation. First, we use an LLM for semantic slot routing to decompose user prompts into character and background descriptions and explicitly select suitable historical references for each component. To isolate the required content, masked memory weaving uses contrasting attention maps conditioned on semantic slots to construct refined semantic masks, selectively exposing relevant tokens from compressed historical KV memories to guide the generation of the current shot. We further introduce coverage adaptive RoPE to adjust temporal offsets and memory retention according to no, partial, or full reference coverage, addressing visual artifacts observed when incomplete historical references are positioned close to the current generation. Extensive experiments demonstrate that Weave Forcing improves cross-shot subject and background consistency while maintaining competitive visual quality and text alignment.
Problem

Research questions and friction points this paper is trying to address.

interactive long video generation
compositional memory routing
cross-shot consistency
visual artifacts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Memory Routing
Masked Memory Weaving
Coverage Adaptive RoPE
Interactive Long Video Generation
Training-free Framework