Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

📅 2025-11-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address challenges in multimodal narrative storytelling—including poor structural control, weak cross-modal consistency, and coarse-grained editing—this paper proposes a graph-node-based multimodal content generation system. Methodologically, it introduces a node-centric editing framework that maps text, images, audio, and video to editable graph nodes; employs a task-selection agent for dynamic generation orchestration; and integrates multimodal large language models, context-aware generation, and natural language understanding to enable node-level precise editing, automatic parallel narrative branching, and cross-modal co-evolution. Contributions include: (1) the first fine-grained, interpretable, human-in-the-loop multimodal narrative iterative generation system; (2) significant performance gains on story outline generation; and (3) user studies demonstrating a 42% improvement in editing efficiency, enhanced creative flexibility, and validated alignment between controllability and effectiveness.

Technology Category

Machine Learning: Multimodal LearningHumans and AI: Game Design — Procedural Content Generation & StorytellingNatural Language Processing: Generation

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: Economic ramifications for generative AI infrastructure and applications
📝 Abstract
We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-language prompts. Each node can integrate text, images, audio, and video, allowing creators to compose multimodal narratives. A task selection agent routes between specialized generative tasks that handle story generation, node structure reasoning, node diagram formatting, and context generation. The interface supports targeted editing of individual nodes, automatic branching for parallel storylines, and node-based iterative refinement. Our results demonstrate that node-based editing supports control over narrative structure and iterative generation of text, images, audio, and video. We report quantitative outcomes on automatic story outline generation and qualitative observations of editing workflows. Finally, we discuss current limitations such as scalability to longer narratives and consistency across multiple nodes, and outline future work toward human-in-the-loop and user-centered creative AI tools.
Problem

Research questions and friction points this paper is trying to address.

Creating multimodal narratives integrating text, audio, images and video
Enabling iterative refinement of story structure through node editing
Maintaining narrative consistency across multiple nodes and modalities
Innovation

Methods, ideas, or system contributions that make the work stand out.

Node-based graph system for multimodal content generation
Task selection agent routes specialized generative tasks
Interface supports targeted editing and automatic branching