🤖 AI Summary
Medieval manuscript images suffer from a scarcity of fine-grained, semantically rich annotations, hindering the application of instance segmentation and multimodal models in cultural heritage.
Method: We propose a zero-shot, model-driven framework for page-level, instance-level annotation of medieval manuscripts. It integrates state-of-the-art instance segmentation, zero-shot vision-language models, and multimodal deep learning to achieve fully automatic region segmentation, content recognition, and semantic description of manuscript pages. Crucially, it systematically uncovers “invisible” graphical and textual elements—such as steganographic text, micro-ornaments, and stylized borders—that elude conventional manual annotation.
Contribution/Results: The resulting dataset features pixel-accurate instance masks with cross-modal semantic alignment (e.g., image–text pairs). It significantly improves downstream models’ accuracy and generalization on rare visual concepts in medieval imagery, establishing a scalable, low-dependency annotation paradigm for medieval image computing.
📝 Abstract
We aim to theorize the medieval manuscript page and its contents more holistically, using state-of-the-art techniques to segment and describe the entire manuscript folio, for the purpose of creating richer training data for computer vision techniques, namely instance segmentation, and multimodal models for medieval-specific visual content.