The Devil is in the Details -- From OCR for Old Church Slavonic to Purely Visual Stemma Reconstruction
This study addresses the low OCR accuracy for Old Church Slavonic manuscripts and the reliance of textual phylogeny reconstruction on transcribed texts by proposing a purely vision-driven, end-to-end approach. Through systematic evaluation of conventional OCR systems, machine learning models, and large language models—including GPT-5 and Gemini3-flash—the method integrates an agent-based architecture with retrieval-augmented generation (RAG) post-processing to substantially improve recognition performance. Notably, it achieves the first fully automated, image-only phylogeny reconstruction pipeline, encompassing glyph extraction, clustering, and distance matrix computation. Experimental results demonstrate a character error rate as low as 2–3% and validate the feasibility and effectiveness of the proposed framework on two medieval manuscript corpora.