Score
Designs and builds systems that convert visual inputs (individual frames, screen captures, or frame sequences) into controlled, context-aware textual narratives and layered summaries, supporting persistent narrations across steps and the construction, maintenance, and modification of narrative variants. Implements and analyzes mechanisms for controlling discourse properties, multi-agent summary synthesis, suppression of redundant output on static content, and generation of searchable, actionable text context, and provides reproducible testbeds to evaluate narrative quality and effects.
This study addresses the lack of systematic narrative-theoretic grounding in current research on automatic story generation and comprehension, noting a pronounced lag in handling nonfictional narratives and multi-level modeling compared to understanding tasks. Drawing on classical narratological frameworks—particularly the distinction between story and discourse levels—the work provides the first systematic review of large language models’ application to narrative tasks, uncovering critical gaps in data diversity, depth of theoretical integration, and task balance. Its primary contribution is a theory-driven, fine-grained evaluation framework that moves beyond monolithic “narrative quality” metrics toward a multidimensional, narratologically informed assessment paradigm. The paper further proposes a comprehensive research roadmap to advance cross-domain narrative analysis and context-aware story generation.
To address challenges in multimodal narrative storytelling—including poor structural control, weak cross-modal consistency, and coarse-grained editing—this paper proposes a graph-node-based multimodal content generation system. Methodologically, it introduces a node-centric editing framework that maps text, images, audio, and video to editable graph nodes; employs a task-selection agent for dynamic generation orchestration; and integrates multimodal large language models, context-aware generation, and natural language understanding to enable node-level precise editing, automatic parallel narrative branching, and cross-modal co-evolution. Contributions include: (1) the first fine-grained, interpretable, human-in-the-loop multimodal narrative iterative generation system; (2) significant performance gains on story outline generation; and (3) user studies demonstrating a 42% improvement in editing efficiency, enhanced creative flexibility, and validated alignment between controllability and effectiveness.
Existing story visualization methods overemphasize visual consistency while neglecting narrative structure and authorial intent, resulting in generated images that fail to accurately convey plot logic and emotional nuance. To address this, we propose a training-free multi-agent framework that enables end-to-end image sequence generation through collaborative agents, jointly optimizing narrative fidelity, semantic consistency, and cross-frame contextual coherence. Our key contributions are: (1) the first narrative-structure-driven hierarchical prompt refinement mechanism, dynamically integrating scene, subject, layout, and other generative elements into a unified workflow; and (2) an integrated pipeline comprising narrative parsing, prompt engineering optimization, diffusion model invocation, and inter-frame semantic alignment. Quantitative and qualitative evaluations demonstrate significant improvements over baselines across multiple narrative fidelity metrics—particularly in visualizing critical plot points and affective details—establishing a novel paradigm for story-driven image generation.
This survey addresses the critical multimodal task of vision-driven story generation, systematically reviewing representative works from 2015 to 2024. Motivated by three key problems—lack of a unified analytical framework, ambiguous task boundaries, and outdated evaluation protocols—we propose, for the first time, a cross-task unifying framework encompassing image/video captioning, visual question answering (VQA), and story generation, clarifying methodological transfer patterns and fundamental distinctions. We critically examine prevalent datasets and metrics (e.g., BLEU, CIDEr, SPICE), exposing their limitations in capturing explainability, controllability, and long-range narrative coherence, and advocate for evaluation reforms aligned with these dimensions. Synthesizing advances in deep learning, multimodal alignment (e.g., CLIP-style architectures), and sequence modeling (e.g., Transformers), we identify current bottlenecks and chart a roadmap toward robust, trustworthy visual storytelling—providing both theoretical foundations and practical guidance for next-generation research.
Current generative AI systems lack fine-grained control over narrative progression in story creation. This work proposes a “narrative keyframe” approach, adapting the concept of animation keyframes to text generation by enabling authors to specify constraints—pertaining to plot, character, or perspective—at designated narrative nodes. The model then automatically generates coherent intermediate content, effectively integrating planning and generation into a unified process. This method achieves a synthesis of high-level controllability and character-centric storytelling. User studies demonstrate that it significantly enhances controllability, transparency, and the sense of creative agency during the generation process.
In long-form narrative video editing, creators face high cognitive load when locating plot points, tracking character motivations, and reassembling distributed events across multi-hour footage; existing transcription- or embedding-based methods lack narrative understanding and fail to support creative decision-making. Method: We propose the first prompt-driven, modular video editing architecture integrating semantic indexing, temporal segmentation, guided memory compression, and cross-granularity narrative fusion to enable interpretable modeling of plot, dialogue, emotion, and context. The system replaces traditional timeline operations with natural language prompts, balancing automation efficiency with editorial control. Contribution/Results: Evaluated on 400+ videos, our method significantly improves editing efficiency while strictly preserving narrative coherence. Professional editors rated it highly in usability and creative efficacy, and user studies demonstrated strong preference over baseline approaches.
This work addresses narrative defocus in long-form audiovisual story generation caused by semantic drift and character inconsistency by proposing the first multi-agent storytelling framework based on closed-loop cognitive coordination. The approach formulates story generation as a constraint satisfaction problem, employing an iterative planning–execution–verification–correction mechanism that integrates explicit machine-executable controls—such as identity preservation, spatial composition, and temporal continuity—with cross-modal feedback to maintain high-level narrative intent over extended durations. Additionally, the study introduces MUSEBench, a reference-free, open-ended evaluation protocol. Experimental results demonstrate that the proposed method significantly outperforms existing baselines in long-horizon narrative coherence, cross-modal identity consistency, and cinematic visual quality.
This work addresses the challenge of long-range inconsistency in multi-frame story illustration generation, where character identity, layout, and emotional expression often drift across frames. To mitigate this, the authors propose the S2ED framework, which employs a multi-agent collaboration mechanism to decompose a full narrative into explicit, editable executable descriptions, enabling coherent narrative segmentation, anchored character attributes, and enhanced spatial-emotional cues. S2ED further introduces a training-free, model-agnostic prompting layer that supports interpretable state propagation and localized editing to correct inter-frame drift. Evaluated on the Flintstones and Shakoo Maku datasets, S2ED outperforms strong prompting baselines, large language model planners, and trainable approaches in both automatic metrics and human assessments, and has been successfully integrated into an end-to-end children’s picture book generation system.
Existing long-form video generation methods lack systematic narrative planning and cross-scene visual consistency, struggling to maintain coherent character and environmental representations across multiple scenes. This work proposes the first multi-agent collaborative framework for long video generation, in which specialized agents jointly negotiate narrative structure, visual continuity, and production quality. The approach integrates a hierarchical narrative engine with a dependency-aware, cross-temporal tracking mechanism for characters and environments, further enhanced by retrieval-augmented generation and vision-language model–guided monitoring optimization. This method substantially improves narrative coherence and cross-scene visual consistency, significantly outperforming current short-clip generation techniques in both story structure and visual fidelity.
Existing visual storytelling approaches predominantly rely on textual inputs alone, struggling to effectively integrate multimodal conditions such as character identity images, scene backgrounds, and shot types, which limits their customizability and cinematic expressiveness. To address this, this work proposes VstoryGen, a novel framework that, for the first time, jointly models these three conditioning modalities within a unified multimodal large language model and introduces an explicit shot-type control mechanism. Leveraging a parameter-efficient prompt-tuning strategy trained on cinematic data, VstoryGen enables customizable visual story generation with high consistency and diversity. Experimental results on two newly established evaluation benchmarks demonstrate that VstoryGen significantly outperforms existing methods in character-scene consistency, image-text alignment, and shot control, thereby enhancing narrative coherence and cinematic language expression.
This work addresses the challenges of structural inconsistency, missing content, and cross-section incoherence commonly encountered in automatically generated scientific papers, particularly between narrative text, experimental evidence, and visual elements. To resolve these issues, the authors propose a multi-agent collaborative framework grounded in a persistent shared visual contract, comprising architect, writer, optimizer, renderer, and evaluator agents. These agents operate within a generate–evaluate–adapt loop that dynamically updates the contract to align textual structure with visual components throughout the document. The approach introduces, for the first time, a contract-driven mechanism to regulate multi-agent collaboration. Evaluated on the Jericho corpus, the method achieves an expert rating of 6.145, significantly outperforming DirectChat (3.963) and Fars (5.197), thereby demonstrating substantial improvements in both structural coherence and text–figure consistency.