Score
Designs and builds comic-format visual narratives that translate datasets into sequential panels, integrating visual encodings, annotations, and explanatory captions to communicate data-driven stories; this includes AI-assisted generation of panel layouts, imagery, and text to automate or augment data-to-comic creation.
Narrative visualization of complex data remains challenging in terms of interpretability and engagement. Method: We conduct a systematic literature review of 66 papers to propose the first end-to-end, four-stage reference model—Analysis, Narrative, Visualization, and Interaction—and distill eight core tasks, including insight extraction and author assistance. We further introduce a unified technical framework integrating large language models, multimodal understanding and generation, data insight mining, and human-AI collaboration, and systematically evaluate performance boundaries and challenges across tasks. Contribution/Results: We construct a structured knowledge graph that clarifies technological applicability scopes and open research questions, delivering an actionable roadmap and evaluation guidelines for researchers and practitioners in narrative visualization powered by foundation models.
Current computational narrative analysis insufficiently addresses comics—a multimodal medium integrating text and images—particularly lacking scene-level narrative arc annotations, thereby hindering advances in multimodal narrative understanding and computational comic analysis. To address this gap, we introduce ComicArc, the first fine-grained, scene-level narrative arc dataset for comics, comprising 154 public-domain comic stories. We propose a text-image cue–guided scene segmentation pipeline, rigorously refined and validated through expert human annotation. ComicArc establishes the first systematic annotation guidelines for scene boundaries and narrative arcs in comics, providing a novel paradigm and benchmark resource for multimodal narrative research. Furthermore, we construct the first comic scene segmentation benchmark, achieving significant improvements in narrative structure recognition performance. This work advances computational comic analysis from page- or panel-level processing toward semantically coherent, scene-level modeling.
This survey addresses the critical multimodal task of vision-driven story generation, systematically reviewing representative works from 2015 to 2024. Motivated by three key problems—lack of a unified analytical framework, ambiguous task boundaries, and outdated evaluation protocols—we propose, for the first time, a cross-task unifying framework encompassing image/video captioning, visual question answering (VQA), and story generation, clarifying methodological transfer patterns and fundamental distinctions. We critically examine prevalent datasets and metrics (e.g., BLEU, CIDEr, SPICE), exposing their limitations in capturing explainability, controllability, and long-range narrative coherence, and advocate for evaluation reforms aligned with these dimensions. Synthesizing advances in deep learning, multimodal alignment (e.g., CLIP-style architectures), and sequence modeling (e.g., Transformers), we identify current bottlenecks and chart a roadmap toward robust, trustworthy visual storytelling—providing both theoretical foundations and practical guidance for next-generation research.
Comic understanding faces core challenges including highly variable artistic styles, nonlinear reading orders, strong visual-textual coupling, and loosely structured narratives. Method: This paper introduces LoCU (Layer of Comics Understanding), the first hierarchical task framework that unifies multi-level objectives—from low-level perception to high-level narrative reasoning—while establishing a structured survey paradigm through systematic analysis of 30+ datasets and methods. Contribution/Results: LoCU enables systematic categorization and critical analysis of existing approaches spanning multimodal vision-language models, document understanding, visual question answering, object detection, and narrative modeling. It identifies key bottlenecks, such as cross-page temporal modeling and style-robust inference, and releases Awesome Comics Understanding—an open-source resource repository—to establish an extensible methodological foundation and collaborative platform for intelligent comic research.
To address the challenge of narrative inaccessibility for visually impaired readers due to comics’ strong visual dependency, this paper introduces MagiV3—the first unified multimodal vision-language model designed for comic understanding and literary narrative generation. Methodologically, it integrates OCR, panel segmentation, character and speech bubble localization, character grounding, and a large vision-language model (VLM) to enable end-to-end generation of coherent, literary text from raw comic panels. Key contributions include: (1) releasing the first high-quality, manually annotated comic panel dataset comprising 3,300+ samples with precise character positions and semantic annotations; (2) proposing a novel collaborative framework that jointly leverages a dedicated visual understanding module and a VLM, markedly improving narrative coherence and literary quality; and (3) demonstrating through extensive experiments that generated narratives significantly outperform baselines in plot completeness, character relationship depiction, and scene atmosphere conveyance—thereby enabling accessible, in-depth reading for visually impaired users.
This paper addresses the challenge of structured semantic understanding in visual narratives (e.g., comics). We propose a hierarchical multimodal knowledge graph framework that decomposes narratives into three granular levels: story arcs, event segments, and panels—unifying semantic and spatiotemporal modeling across levels. Our key innovation is a novel multi-granularity alignment mechanism enabling panel-level visual–textual coupling and cross-level symbolic reasoning. The framework constructs multimodal graphs, fuses them hierarchically, and adapts to a manually annotated Manga109 subset. Evaluated on four tasks—action retrieval, dialogue tracking, character mapping, and panel temporal reconstruction—it achieves high precision and recall. Experiments demonstrate significant advantages in interpretability, consistent multimodal representation, and cross-task generalization.
This work addresses the challenge that existing infographics authoring tools struggle to maintain narrative coherence and align with users’ storytelling goals during the design process. To bridge this gap, the authors propose a narrative-centered, three-stage human-AI collaborative framework—encompassing story construction, visual encoding, and spatial layout—and implement it in a system called InfoAlign. By integrating natural language processing, semantic alignment–based design recommendations, and parametric layout generation, InfoAlign enables the transformation of unstructured text into a coherent narrative, recommends visually consistent encodings, and produces layout blueprints—all while preserving full user control throughout the workflow. InfoAlign represents the first approach to explicitly align story semantics with visual expression, significantly enhancing both narrative coherence and the collaborative experience in data storytelling.
This work addresses the challenges of automatically generating animated data videos that cohesively integrate dynamic visualizations with synchronized narration—a task requiring careful coordination of visual encoding, temporal progression, and narrative structure, yet lacking a standardized evaluation benchmark. To bridge this gap, we introduce DataReel, the first benchmark specifically designed for animated data video storytelling, comprising 328 real-world video stories. We further propose a large language model–based multi-agent framework that decomposes the generation process into planning, generation, and validation stages, emulating human-like narrative workflows to enable end-to-end collaborative creation. Experimental results demonstrate that our approach significantly outperforms direct prompting baselines in both automatic and human evaluations, uncovering key mechanisms and challenges in the synergistic interplay among animation, narration, and visual emphasis.
This study addresses the challenge that students often struggle to effectively interpret traditional data visualizations due to limited visual literacy, while high-quality data comics—though potentially more accessible—are costly to produce, hindering their educational adoption. To bridge this gap, the authors propose leveraging generative artificial intelligence (GenAI) to support the creation of data comics and evaluate their pedagogical efficacy through a controlled experiment focused on information retrieval and insight comprehension tasks. Findings indicate that students using AI-generated data comics outperformed those using conventional visualizations and perceived the comics as more engaging and comprehensible. This work provides the first empirical evidence of the educational effectiveness of GenAI-assisted data comics and systematically examines associated ethical considerations, including accuracy, potential for misinterpretation, and copyright implications.
This work addresses the challenge of long-range inconsistency in multi-frame story illustration generation, where character identity, layout, and emotional expression often drift across frames. To mitigate this, the authors propose the S2ED framework, which employs a multi-agent collaboration mechanism to decompose a full narrative into explicit, editable executable descriptions, enabling coherent narrative segmentation, anchored character attributes, and enhanced spatial-emotional cues. S2ED further introduces a training-free, model-agnostic prompting layer that supports interpretable state propagation and localized editing to correct inter-frame drift. Evaluated on the Flintstones and Shakoo Maku datasets, S2ED outperforms strong prompting baselines, large language model planners, and trainable approaches in both automatic metrics and human assessments, and has been successfully integrated into an end-to-end children’s picture book generation system.
Existing end-to-end generative models struggle to precisely control layout geometry, visual references, and cross-panel consistency in comic generation. This work proposes an agent-based, multi-stage generation framework that decouples the creative pipeline into modular stages—story planning, character-scene anchoring, layout construction, reference-guided rendering, page composition, and text typesetting. By incorporating a story-paragraph memory mechanism and explicit intermediate representations, the approach enables editable control over layout, visual assets, and textual elements. The method significantly outperforms end-to-end baselines in layout fidelity, cross-panel consistency, and overall generation quality, while supporting flexible human intervention. This provides a controllable and efficient solution for generating long-form comics.