🤖 AI Summary
This study addresses the limitation of existing multimodal generative models, which produce individually plausible outputs across modalities yet lack cross-modal narrative coherence. To investigate this, we construct a story-based multimodal benchmark that evaluates the synergistic continuation capabilities of images, narration, and speech. Furthermore, we propose a consistency-centric LLM-based evaluation metric to quantify cross-modal context preservation, and adopt a modeling paradigm that integrates VLM-guided planning with native Any-to-Any generation. Experimental results demonstrate that strong VLM orchestration strategies yield the most reliable performance, while revealing that visual continuity and semantic completeness remain the primary bottlenecks in current multimodal narrative generation.
📝 Abstract
Omnimodal evaluation should go beyond independent text, image, and speech production: individually plausible outputs may not express a coherent shared event. We introduce Omni-StoryBench, a story-grounded omnimodal benchmark evaluating whether models can coherently continue stories across image, narration, and speech. Each instance provides a current storybook page and structured next-page conditions, requiring models to generate the next illustration, narration, and spoken character utterance. Omni-StoryBench contains 900 rigorously validated story transitions from openly licensed children's books, with ground-truth next-page references and speech metadata. We evaluate systems with modality-specific metrics and consistency-centered LLM-as-a-judge rubrics for context preservation, condition following, reference consistency, and cross-modal coherence. Across 32 baseline configurations spanning orchestration, semi-orchestration, and native any-to-any paradigms, we find orchestration with strong VLM planning most reliable, while current native omnimodal models often struggle with output completeness and controllability. Our analysis shows text-side performance is associated with image and speech quality, but image generation and visual continuity form the clearest observed bottleneck among the evaluated configurations. These results position Omni-StoryBench as a system-level benchmark measuring coherent omnimodal generation beyond isolated modality quality.