From Panels to Prose: Generating Literary Narratives from Comics

📅 2025-03-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of narrative inaccessibility for visually impaired readers due to comics’ strong visual dependency, this paper introduces MagiV3—the first unified multimodal vision-language model designed for comic understanding and literary narrative generation. Methodologically, it integrates OCR, panel segmentation, character and speech bubble localization, character grounding, and a large vision-language model (VLM) to enable end-to-end generation of coherent, literary text from raw comic panels. Key contributions include: (1) releasing the first high-quality, manually annotated comic panel dataset comprising 3,300+ samples with precise character positions and semantic annotations; (2) proposing a novel collaborative framework that jointly leverages a dedicated visual understanding module and a VLM, markedly improving narrative coherence and literary quality; and (3) demonstrating through extensive experiments that generated narratives significantly outperform baselines in plot completeness, character relationship depiction, and scene atmosphere conveyance—thereby enabling accessible, in-depth reading for visually impaired users.

Technology Category

Computer Vision: Language and VisionNatural Language Processing: Language Grounding & Multi-modal NLPMachine Learning: Large Multimodal Models (LMMs)

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsWeb Mining and Content Analysis: Large pretrained models with web dataSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Comics have long been a popular form of storytelling, offering visually engaging narratives that captivate audiences worldwide. However, the visual nature of comics presents a significant barrier for visually impaired readers, limiting their access to these engaging stories. In this work, we provide a pragmatic solution to this accessibility challenge by developing an automated system that generates text-based literary narratives from manga comics. Our approach aims to create an evocative and immersive prose that not only conveys the original narrative but also captures the depth and complexity of characters, their interactions, and the vivid settings in which they reside. To this end we make the following contributions: (1) We present a unified model, Magiv3, that excels at various functional tasks pertaining to comic understanding, such as localising panels, characters, texts, and speech-bubble tails, performing OCR, grounding characters etc. (2) We release human-annotated captions for over 3300 Japanese comic panels, along with character grounding annotations, and benchmark large vision-language models in their ability to understand comic images. (3) Finally, we demonstrate how integrating large vision-language models with Magiv3, can generate seamless literary narratives that allows visually impaired audiences to engage with the depth and richness of comic storytelling.
Problem

Research questions and friction points this paper is trying to address.

Generating text narratives from comics for visually impaired readers
Automating comic understanding with a unified model (Magiv3)
Creating immersive prose that captures comic depth and complexity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Automated system generates text from manga comics
Magiv3 model localizes panels, characters, and texts
Integrates vision-language models for seamless narratives
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
R
Ragav Sachdeva
Visual Geometry Group, Dept. of Engineering Science, University of Oxford
Andrew Zisserman
Andrew Zisserman
University of Oxford
Computer VisionMachine Learning