🤖 AI Summary
This study addresses the narrative discontinuity and visual drift in multi-shot video generation caused by the absence of explicit state propagation. To this end, we propose a state-grounded agent framework that decouples semantic planning from visual observation. By employing an event-driven mechanism to maintain a structured world state, our approach effectively prevents error accumulation. Furthermore, it achieves precise long-video rendering through canonical reference construction, constraint compilation, and a bounded evaluation-guided repair loop. Extensive experiments on a self-constructed benchmark demonstrate that the proposed method comprehensively outperforms existing state-of-the-art approaches in terms of narrative quality, temporal coherence, and visual consistency.
📝 Abstract
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.