What Comes Next? Omni-StoryBench for Evaluating Story-Grounded Omnimodal Generation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing multimodal generative models, which produce individually plausible outputs across modalities yet lack cross-modal narrative coherence. To investigate this, we construct a story-based multimodal benchmark that evaluates the synergistic continuation capabilities of images, narration, and speech. Furthermore, we propose a consistency-centric LLM-based evaluation metric to quantify cross-modal context preservation, and adopt a modeling paradigm that integrates VLM-guided planning with native Any-to-Any generation. Experimental results demonstrate that strong VLM orchestration strategies yield the most reliable performance, while revealing that visual continuity and semantic completeness remain the primary bottlenecks in current multimodal narrative generation.
📝 Abstract
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.
Problem

Research questions and friction points this paper is trying to address.

omnimodal generation
story-grounded evaluation
cross-modal coherence
benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Omnimodal Generation
Story-Grounded Benchmark
Cross-Modal Coherence
LLM-as-a-Judge
Visual Continuity
🔎 Similar Papers
No similar papers found.