The Devil is in the Details -- From OCR for Old Church Slavonic to Purely Visual Stemma Reconstruction

📅 2026-04-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the low OCR accuracy for Old Church Slavonic manuscripts and the reliance of textual phylogeny reconstruction on transcribed texts by proposing a purely vision-driven, end-to-end approach. Through systematic evaluation of conventional OCR systems, machine learning models, and large language models—including GPT-5 and Gemini3-flash—the method integrates an agent-based architecture with retrieval-augmented generation (RAG) post-processing to substantially improve recognition performance. Notably, it achieves the first fully automated, image-only phylogeny reconstruction pipeline, encompassing glyph extraction, clustering, and distance matrix computation. Experimental results demonstrate a character error rate as low as 2–3% and validate the feasibility and effectiveness of the proposed framework on two medieval manuscript corpora.

Technology Category

Computer Vision: Low Level & Physics-based VisionCognitive Modeling & Cognitive Systems: Agent ArchitecturesNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphsEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
The age of artificial intelligence has brought many new possibilities and pitfalls in many fields and tasks. The devil is in the details, and those come to the fore when building new pipelines and executing small practical experiments. OCR and stemmatology are no exception. The current investigation starts comparing a range of OCR-systems, from classical over machine learning to LLMs, for roughly 6,000 characters of late handwritten church slavonic manuscripts from the 18th century. Focussing on basic letter correctness, more than 10 CS OCR-systems among which 2 LLMs (GPT5 and Gemini3-flash) are being compared. Then, post-processing via LLMs is assessed and finally, different agentic OCR architectures (specialized post-processing agents, an agentic pipeline and RAG) are tested. With new technology elaborated, experiments suggest, church slavonic CER for basic letters may reach as low as 2-3% but elaborated diacritics could still present a problem. How well OCR can prime stemmatology as a downstream task is the entry point to the second part of the article which introduces a new stemmatic method based solely on image processing. Here, a pipeline of automated visual glyph extraction, clustering and pairwise statistical comparison leading to a distance matrix and ultimately a stemma, is being presented and applied to two small corpora, one for the church slavonic Gospel of Mark from the 14th to 16th centuries, one for the Roman de la Rose in French from the 14th and 15th centuries. Basic functioning of the method can be demonstrated.
Problem

Research questions and friction points this paper is trying to address.

OCR
Church Slavonic
Stemmatology
Manuscript Analysis
Visual Stemma Reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

visual stemma reconstruction
agentic OCR
glyph clustering
historical manuscript OCR
RAG-based post-processing
🔎 Similar Papers
No similar papers found.