manuscript ocr

Designs and implements systems and pipelines that convert images of manuscripts and other handwritten or historical documents into machine-readable transcriptions, including end-to-end OCR/HTR models, lightweight inference architectures, and interactive transcription tools. Work includes OCR error analysis and correction, modelling scribal shorthand and ligatures, handling degraded or low-quality input, and post-processing to reduce character- and token-level error rates.

manuscriptocr

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes an end-to-end approach integrating layout analysis, text line detection, and optical character recognition (OCR) to address the challenge of preserving special characters and symbols in full-page transcriptions of Latin historical documents from the 15th to 16th centuries. By incorporating masked autoencoder technology, the method uniformly handles handwritten, printed, and multilingual mixed texts, accurately extracting text lines while fully retaining original typographic styles and semantic symbols. Experimental results on multiple historical document datasets demonstrate that the proposed framework achieves high accuracy and efficiency in digitizing complex historical pages, significantly outperforming existing state-of-the-art methods.

historical document transcriptionLatin scriptlayout analysis

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

May 26, 2025
SG
Shuhao Guan
🏛️ University College Dublin | Trinity College Dublin | University of Toronto | Shanghai University

To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.

Combining image restoration with OCR error correctionImproving text extraction from degraded historical documentsReducing character error rates in multilingual historical texts

This study addresses the challenges of optical character recognition (OCR) for eighteenth-century printed texts, which include archaic glyphs, non-standardized spelling, and print degradation—factors that render conventional evaluation metrics inadequate for assessing scholarly reliability. The authors propose a combined approach using length-weighted accuracy and hypothesis-driven error analysis to systematically compare the performance of the specialized OCR model TrOCR against the general-purpose vision-language model Qwen on historical English text lines. Their findings reveal systematic differences rooted in architectural inductive biases: Qwen exhibits lower overall error rates and greater robustness to degraded inputs but implicitly normalizes spelling, whereas TrOCR more faithfully preserves original orthography at the cost of susceptibility to cascading errors. The work underscores the necessity of aligning model selection with scholarly risk assessment in the digitization of historical documents.

Character Error RateError PatternsHistorical OCR

Medieval Latin legal manuscripts pose significant transcription challenges due to pervasive abbreviations, degraded paleographic features, and complex juridical context, resulting in high error rates. To address this, we propose a four-stage hybrid workflow: (1) initial transcription via a domain-specific handwritten text recognition (HTR) model; (2) image-text joint post-correction using multimodal large language models (MLLMs); (3) prompt-engineered Latin abbreviation expansion; and (4) integration of a named entity consolidation (NEC) module. A key contribution is the construction of “clean ground-truth” training data, substantially enhancing model capacity for medieval linguistic patterns and legal terminology. Evaluated on an academic-grade ground-truth benchmark, our method achieves word error rates (WER) of 2%–7%. Case studies confirm that outputs preserve semantic fidelity while yielding structured, analysis-ready text suitable for historical semantics and digital humanities research—thereby markedly improving both transcription accuracy and scholarly utility of medieval legal documents.

Automating laborious transcription stages while maintaining scholarly qualityCorrecting and expanding abbreviated text using multimodal LLM techniquesTranscribing challenging medieval Latin legal documents accurately

This study addresses the challenges of machine translation for medieval Latin manuscripts, including low-resource transcription, archaic language recognition, and image noise. The authors introduce the first end-to-end image-to-translation evaluation framework and release the Interpres-Parallel-Corpus, a dataset comprising 1,383 aligned lines of manuscript images, transcriptions, and expert translations. Through systematic comparison of domain-specific OCR, vision-language models (VLMs), retrieval-augmented generation (RAG), and post-OCR correction strategies, they find that a minimal OCR+VLM pipeline yields the highest translation quality, whereas more complex multi-component architectures suffer from error propagation and prompt saturation. Experiments reveal that domain-specific OCR reduces character error rates by up to 4.3× compared to general-purpose VLMs, highlighting a “complexity paradox” wherein simplified pipelines prove more effective.

historical documentslow-resourcemachine translation

Latest Papers

What's happening recently
View more

This work addresses the challenge of automatically deciphering historical encrypted manuscripts, where conventional two-stage pipelines often fail due to transcription errors that propagate into the decryption stage. To overcome this limitation, we propose the first end-to-end image-to-plaintext decryption framework that bypasses intermediate transcription entirely, jointly modeling computer vision and sequence decoding within a unified architecture to prevent error accumulation. To facilitate training, we develop a large-scale synthetic data generation pipeline producing cipher-like sequences grounded in realistic visual and linguistic priors. Experiments on the Copiale cipher demonstrate that our approach significantly outperforms traditional two-stage methods, confirming the efficacy of end-to-end modeling in enhancing both robustness and overall decipherment performance.

automatic deciphermentencrypted handwritten documentshistorical manuscripts

This study addresses a critical yet overlooked issue in historical document digitization: while existing OCR systems prioritize low character error rates (CER), they often neglect semantic hallucinations—such as altered named entities and contextual substitutions—that compromise content fidelity. Focusing on Uruguayan dictatorship-era microfilm documents, this work systematically reveals that vision-language models (VLMs), despite outperforming traditional OCR in CER and word error rate (WER), frequently introduce semantic distortions including orthographic normalization, fabricated content, and substitution of key entities. Through a combination of quantitative metrics and qualitative human analysis, the research demonstrates that conventional evaluation paradigms fail to capture these subtle but significant inaccuracies, thereby advocating for a new semantic trustworthiness assessment framework that transcends character-level accuracy.

hallucinationhistorical documentsOCR

This study addresses the challenge that existing handwritten text recognition methods lack interpretable visual metrics suitable for paleographic analysis. The authors propose a novel architecture requiring only line-level transcription supervision, integrating a Transformer-based detection model, prototype-driven character representation learning, and a line-level reconstruction module to enable weakly supervised character localization and deformation modeling. This approach is the first to support automated paleographic measurements of individual characters, bigrams, and inter-character spacing. Evaluated on 160 pages from the 14th-century manuscript BnF fr. 2813, the method effectively distinguishes scribal styles and reveals subtle writing variations using only single-column text, significantly outperforming baselines such as Learnable Typewriter. Code and data are publicly released.

handwritten text recognitionhistorical scriptmorphological analysis

Hot Scholars

LJ

Lianwen Jin

Professor of Electronic and Information Engineering, South China University of Technology
Optical Character Recognition (OCR)Computer VisionDocument AIMultimodal LLMs
XB

Xiang Bai

Huazhong University of Science and Technology (HUST)
Computer VisionOCR
NJ

Nevidu Jayatilleke

University of Moratuwa, Sri Lanka
Computational LinguisticsArtificial IntelligenceMachine Learning
JT

Jingqun Tang

ByteDance Inc.
Computer VisionDocument IntelligenceMLLMMultimodal Generative Models
MP

Marco Peer

HEIA-FR
Document AnalysisComputer VisionMachine Learning