🤖 AI Summary
This study addresses the issue of plot inconsistency in long-form novel generation by large language models due to the lack of explicit memory, proposing the NarraWorld framework. This approach reformulates memory construction as world modeling, integrating multi-granular narrative information through structured graphs. It introduces a novel four-view connection model and a hierarchical aggregation atomic closure mechanism to explicitly recover implicit dependencies. Furthermore, an evidence graph-based retrieval strategy with planning reconstruction is designed to efficiently assemble context under token budget constraints. Experimental results demonstrate that NarraWorld achieves state-of-the-art performance across three writing benchmarks and successfully transfers to role-playing and general long-term memory tasks.
📝 Abstract
LLM agents that write long-form fiction need an explicit memory of the evolving storyworld to keep new events consistent with established facts. Such memory must keep heterogeneous narrative information distinct, integrate story developments across granularities, and recover dependencies that a writing request leaves implicit. We present NarraWorld, a structured memory system for long-form writing that treats memory construction as world creation. From a shared evidence-grounded graph, NarraWorld derives four connected views: world facts, per-character beliefs, open developments, and hypothetical branches (possible-world continuations). Hierarchical aggregation with atomic closure consolidates events into scenes, plotlines, and plots, keeping each higher-level node traceable to its constituent source spans. For retrieval, planned reconstruction infers a query's dependencies from the current narrative situation and a preview of memory, then assembles the relevant records within a token budget. Across three writing benchmarks, NarraWorld achieves the strongest aggregate results. Its memory also transfers to situated role-playing and largely preserves recall on a general-purpose long-term memory benchmark, paving the way for agents that sustain coherent storyworlds across diverse narrative tasks.