Score
Designs and implements systems and pipelines that convert images of manuscripts and other handwritten or historical documents into machine-readable transcriptions, including end-to-end OCR/HTR models, lightweight inference architectures, and interactive transcription tools. Work includes OCR error analysis and correction, modelling scribal shorthand and ligatures, handling degraded or low-quality input, and post-processing to reduce character- and token-level error rates.
This survey addresses the robustness challenges in Handwritten Text Recognition (HTR) arising from high variability in handwriting styles, document layouts, and image quality. It systematically traces the evolution from heuristic approaches to end-to-end, document-level deep learning models. We propose the first unified analytical framework for HTR—comprehensively covering methodologies, benchmark evaluations, mainstream datasets, and performance comparisons—while explicitly defining recognition granularities: line-level and supra-line-level (i.e., paragraph- or document-level). The framework integrates CNNs, RNNs, Transformers, and attention mechanisms, synergizing sequence modeling strategies such as Connectionist Temporal Classification (CTC) and attention-based Seq2Seq to support multi-granularity recognition. Synthesizing over 100 works, we clarify the technical trajectory and identify core open challenges: cross-domain generalization, model interpretability, and system-level robustness. This work establishes a foundational theoretical and practical reference for developing real-world document understanding systems.
This work proposes an end-to-end approach integrating layout analysis, text line detection, and optical character recognition (OCR) to address the challenge of preserving special characters and symbols in full-page transcriptions of Latin historical documents from the 15th to 16th centuries. By incorporating masked autoencoder technology, the method uniformly handles handwritten, printed, and multilingual mixed texts, accurately extracting text lines while fully retaining original typographic styles and semantic symbols. Experimental results on multiple historical document datasets demonstrate that the proposed framework achieves high accuracy and efficiency in digitizing complex historical pages, significantly outperforming existing state-of-the-art methods.
To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.
This study addresses the challenges of optical character recognition (OCR) for eighteenth-century printed texts, which include archaic glyphs, non-standardized spelling, and print degradation—factors that render conventional evaluation metrics inadequate for assessing scholarly reliability. The authors propose a combined approach using length-weighted accuracy and hypothesis-driven error analysis to systematically compare the performance of the specialized OCR model TrOCR against the general-purpose vision-language model Qwen on historical English text lines. Their findings reveal systematic differences rooted in architectural inductive biases: Qwen exhibits lower overall error rates and greater robustness to degraded inputs but implicitly normalizes spelling, whereas TrOCR more faithfully preserves original orthography at the cost of susceptibility to cascading errors. The work underscores the necessity of aligning model selection with scholarly risk assessment in the digitization of historical documents.
Medieval Latin legal manuscripts pose significant transcription challenges due to pervasive abbreviations, degraded paleographic features, and complex juridical context, resulting in high error rates. To address this, we propose a four-stage hybrid workflow: (1) initial transcription via a domain-specific handwritten text recognition (HTR) model; (2) image-text joint post-correction using multimodal large language models (MLLMs); (3) prompt-engineered Latin abbreviation expansion; and (4) integration of a named entity consolidation (NEC) module. A key contribution is the construction of “clean ground-truth” training data, substantially enhancing model capacity for medieval linguistic patterns and legal terminology. Evaluated on an academic-grade ground-truth benchmark, our method achieves word error rates (WER) of 2%–7%. Case studies confirm that outputs preserve semantic fidelity while yielding structured, analysis-ready text suitable for historical semantics and digital humanities research—thereby markedly improving both transcription accuracy and scholarly utility of medieval legal documents.
This study addresses the challenges of machine translation for medieval Latin manuscripts, including low-resource transcription, archaic language recognition, and image noise. The authors introduce the first end-to-end image-to-translation evaluation framework and release the Interpres-Parallel-Corpus, a dataset comprising 1,383 aligned lines of manuscript images, transcriptions, and expert translations. Through systematic comparison of domain-specific OCR, vision-language models (VLMs), retrieval-augmented generation (RAG), and post-OCR correction strategies, they find that a minimal OCR+VLM pipeline yields the highest translation quality, whereas more complex multi-component architectures suffer from error propagation and prompt saturation. Experiments reveal that domain-specific OCR reduces character error rates by up to 4.3× compared to general-purpose VLMs, highlighting a “complexity paradox” wherein simplified pipelines prove more effective.
为解决复杂历史梵文手稿的OCR难题,提出一种可迭代微调的传统OCR管道,适应目标手稿布局和外观,减少人工标注工作。
This work addresses the challenge of automatically deciphering historical encrypted manuscripts, where conventional two-stage pipelines often fail due to transcription errors that propagate into the decryption stage. To overcome this limitation, we propose the first end-to-end image-to-plaintext decryption framework that bypasses intermediate transcription entirely, jointly modeling computer vision and sequence decoding within a unified architecture to prevent error accumulation. To facilitate training, we develop a large-scale synthetic data generation pipeline producing cipher-like sequences grounded in realistic visual and linguistic priors. Experiments on the Copiale cipher demonstrate that our approach significantly outperforms traditional two-stage methods, confirming the efficacy of end-to-end modeling in enhancing both robustness and overall decipherment performance.
This study addresses a critical yet overlooked issue in historical document digitization: while existing OCR systems prioritize low character error rates (CER), they often neglect semantic hallucinations—such as altered named entities and contextual substitutions—that compromise content fidelity. Focusing on Uruguayan dictatorship-era microfilm documents, this work systematically reveals that vision-language models (VLMs), despite outperforming traditional OCR in CER and word error rate (WER), frequently introduce semantic distortions including orthographic normalization, fabricated content, and substitution of key entities. Through a combination of quantitative metrics and qualitative human analysis, the research demonstrates that conventional evaluation paradigms fail to capture these subtle but significant inaccuracies, thereby advocating for a new semantic trustworthiness assessment framework that transcends character-level accuracy.
研究评估了历史文献数字化流程,采用直接提取、大语言模型后校正及分块提取法,解决了OCR错误传播问题,提高信息检索准确性。
This study addresses the challenge that existing handwritten text recognition methods lack interpretable visual metrics suitable for paleographic analysis. The authors propose a novel architecture requiring only line-level transcription supervision, integrating a Transformer-based detection model, prototype-driven character representation learning, and a line-level reconstruction module to enable weakly supervised character localization and deformation modeling. This approach is the first to support automated paleographic measurements of individual characters, bigrams, and inter-character spacing. Evaluated on 160 pages from the 14th-century manuscript BnF fr. 2813, the method effectively distinguishes scribal styles and reveals subtle writing variations using only single-column text, significantly outperforming baselines such as Learnable Typewriter. Code and data are publicly released.