FP-THD: Full page transcription of historical documents

📅 2026-01-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes an end-to-end approach integrating layout analysis, text line detection, and optical character recognition (OCR) to address the challenge of preserving special characters and symbols in full-page transcriptions of Latin historical documents from the 15th to 16th centuries. By incorporating masked autoencoder technology, the method uniformly handles handwritten, printed, and multilingual mixed texts, accurately extracting text lines while fully retaining original typographic styles and semantic symbols. Experimental results on multiple historical document datasets demonstrate that the proposed framework achieves high accuracy and efficiency in digitizing complex historical pages, significantly outperforming existing state-of-the-art methods.

Technology Category

Natural Language Processing: Sentence-level Semantics, Textual Inference, etc.Machine Learning: Large Multimodal Models (LMMs)Search and Optimization: Mixed Discrete/Continuous Search

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchWeb Mining and Content Analysis: Normalization, clustering, classification, and summarization of Web textGraph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphs
📝 Abstract
The transcription of historical documents written in Latin in XV and XVI centuries has special challenges as it must maintain the characters and special symbols that have distinct meanings to ensure that historical texts retain their original style and significance. This work proposes a pipeline for the transcription of historical documents preserving these special features. We propose to extend an existing text line recognition method with a layout analysis model. We analyze historical text images using a layout analysis model to extract text lines, which are then processed by an OCR model to generate a fully digitized page. We showed that our pipeline facilitates the processing of the page and produces an efficient result. We evaluated our approach on multiple datasets and demonstrate that the masked autoencoder effectively processes different types of text, including handwritten, printed and multi-language.
Problem

Research questions and friction points this paper is trying to address.

historical document transcription
Latin script
special characters
layout analysis
OCR
Innovation

Methods, ideas, or system contributions that make the work stand out.

historical document transcription
layout analysis
masked autoencoder
OCR
full-page digitization
💼 Related Jobs
No related jobs found.
H
H. Neji
J
J. Nogueras-Iso
J
J. Lacasta
M
M'A Latre
F
FJ Garc'ia-Marco