SubjectAnchor: Subject-Aware Memory-to-Video for Multi-Shot Storytelling

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of maintaining cross-shot subject identity and scene consistency in multi-shot video generation by proposing the SubjectAnchor paradigm. Built upon Wan2.2-I2V, this method constructs an explicit visual memory that retrieves keyframes and encodes them as conditional inputs to guide generation. It introduces a novel subject-aware memory-to-video framework incorporating a negative temporal slot separation mechanism, subject-aware temporal rotary position encoding, and memory-aware attention partitioning to effectively suppress cross-shot identity interference. Experimental results demonstrate that the proposed paradigm significantly enhances subject consistency while preserving high visual quality and precise controllability over shot-by-shot prompts.
📝 Abstract
We present SubjectAnchor, a Subject-Aware Memory-to-Video paradigm for multi-shot storytelling in which the current shot is generated by conditioning on explicit visual memories extracted from previous shots. The objective is to preserve subject identity and scene consistency across cuts while retaining the controllability of shot-wise prompting. Built on Wan2.2-I2V-A14B, SubjectAnchor contains three key components: subject-related memory construction, subject-aware temporal rotary position encoding, and memory-aware attention partition. For each target shot, the method constructs a compact memory bank by tracing each required subject to its historical appearance and retrieving the most relevant precomputed keyframes. These memory frames are encoded into the model input as explicit visual conditions, while different subjects are assigned to separated negative temporal slots to reduce identity interference. In addition, memory-aware attention partition regulates the interaction between memory tokens and generated content within a shared backbone. This formulation preserves the appearance anchoring of explicit visual memory while remaining compatible with script-driven shot-by-shot generation. Experiments show that SubjectAnchor improves cross-shot identity consistency over representative memory-based and holistic baselines while maintaining competitive visual quality.
Problem

Research questions and friction points this paper is trying to address.

multi-shot storytelling
subject identity consistency
video generation
scene consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Subject-Aware Memory-to-Video
Multi-Shot Storytelling
Temporal Rotary Position Encoding
Memory-Aware Attention Partition
Cross-Shot Identity Consistency
🔎 Similar Papers
2024-05-22Annual Meeting of the Association for Computational LinguisticsCitations: 2