π€ AI Summary
This study addresses the challenge of retrieving large-scale digitized collections of historical manuscript illustrations, which lack semantic classification. To this end, we construct the first large-scale, fine-grained classification benchmark dataset comprising 15,000 images across 22 categories. Through a systematic evaluation of models including ConvNeXt, Vision Transformer, CLIP, and XGBoost on this task, we reveal the limitations of general-purpose foundation models in specialized historical visual domains. Experimental results demonstrate that a specifically fine-tuned ConvNeXt classifier achieves optimal performance with an accuracy of 88.9%, confirming that domain-specific fine-tuning strategies significantly outperform zero-shot general-purpose models. This work provides effective methodological support for digital cultural heritage research.
π Abstract
Historical manuscript illustrations preserve rich visual evidence of past cultures. They depict people, animals, plants, diagrams, music notations, and decorative forms. Although large digitization projects have made many manuscripts available online, the material itself remains difficult to explore at scale. Extraction systems can find illustrations on manuscript pages, but without meaningful categories, large collections remain hard to search and explore. We address this gap by introducing a manually labeled dataset of 15,000 illustrations from manuscripts dating back hundreds of years across 22 categories, and evaluating modern vision models for image classification on this task. The problem is challenging due to stylistic diversity, degradation, and semantic ambiguity, with many images that fit more than one category. We compare fine-tuned CNN and Transformer-based classifiers, zero-shot CLIP, embedding-based classifiers, and direct vision-language models. Results show that fine-tuned image classifiers perform best overall, with ConvNeXt reaching 88.9% accuracy and 81.3% macro-F1. Using CLIP embeddings with XGBoost provides a strong alternative. In contrast, zero-shot CLIP and direct vision-language classification perform substantially worse, highlighting the limits of general-purpose models in this domain. Beyond overall performance, the analysis reveals which categories are visually separable and where errors reflect genuine semantic overlap, suggesting that some limitations arise from the taxonomy itself.