optical character recognition

Designs, implements, or evaluates systems that detect, segment, and convert printed or handwritten text in images or scanned documents into machine-encoded characters; this includes preprocessing (deskewing, denoising, layout analysis), recognition models (character/sequence classifiers, OCR engines), and postprocessing (language-model correction, formatting) to produce structured textual output.

opticalcharacterrecognition

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.63
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$180K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Low-quality invoice images—characterized by complex table structures, severe noise, and heterogeneous layouts—significantly degrade OCR accuracy. To address this, we propose an end-to-end OCR-driven pipeline for tabular data extraction. Our method introduces a dynamic image preprocessing mechanism to enhance readability of degraded invoices; designs an adaptive table boundary detection and row-column mapping algorithm to robustly localize non-standard tables and semantically align cells; and integrates Tesseract OCR with customized post-processing logic for accurate text recognition and structured reconstruction. Experiments on a real-world invoice dataset demonstrate substantial improvements: +12.7% in field-level accuracy and enhanced layout consistency. The pipeline enables high-precision financial automation and digital archival, exhibiting strong engineering deployability in production environments.

Automating financial workflows via invoice digitizationExtracting structured tabular data from noisy invoicesImproving accuracy in OCR-based table recognition

This work addresses the challenge of accurately and efficiently digitizing complex documents containing handwritten content, irregular tables, and heterogeneous layouts—tasks that remain difficult for conventional OCR systems and current large language models. The authors propose an interactive document digitization system that integrates layout-aware parsing, OCR, and a large language model, enhanced by a user-in-the-loop correction propagation mechanism. Leveraging layout-aware inference, the system automatically generalizes user edits or natural language instructions applied to a local region to structurally similar regions across the document. In a user study (n=12), this approach significantly improved correction efficiency, reduced repetitive manual operations, and enabled more controllable and effective reconstruction of document structure and content.

document digitizationhandwritten contentheterogeneous layouts

Digitization of Document and Information Extraction using OCR

Jun 11, 2025
RS
Rasha Sinha
🏛️ RV College of Engineering

To address the limited flexibility and weak semantic understanding in information extraction from mixed scanned images and natively digital documents, this paper proposes a two-stage OCR-LLM collaborative framework. In the first stage, multi-engine OCR combined with layout-aware parsing (e.g., PDFMiner/LayoutParser) performs preliminary text and structural extraction. In the second stage, a fine-tuned large language model (LLM), enhanced by context-aware prompt engineering, achieves cross-format, multi-layout semantic parsing of key fields and generates confidence scores. The work introduces the novel “OCR pre-extraction + LLM post-parsing” paradigm, integrating visual layout cues with linguistic context to resolve format-induced ambiguities. Experiments demonstrate a 32% improvement in key-field extraction accuracy, support for over ten document types, a 27% reduction in average latency, and significantly enhanced layout robustness and generalization capability.

Combining OCR and LLMs for structured, context-aware data extractionEvaluating OCR tools for accuracy, layout recognition, and speedExtracting accurate text from mixed scanned and digital documents

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

May 26, 2025
SG
Shuhao Guan
🏛️ University College Dublin | Trinity College Dublin | University of Toronto | Shanghai University

To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.

Combining image restoration with OCR error correctionImproving text extraction from degraded historical documentsReducing character error rates in multilingual historical texts

General Detection-based Text Line Recognition

Sep 25, 2024
RB
Raphael Baena
🏛️ Ecole des Ponts | Univ Gustave Eiffel | CNRS

Addressing the challenges of cross-script generalization (e.g., Latin, Chinese, cipher scripts) and high annotation cost for character-level supervision in text-line recognition (OCR/HTR), this paper proposes DTLR—an end-to-end detection-based text-line recognition framework. DTLR reformulates text-line recognition as a parallel character detection task, abandoning autoregressive decoding. It is trained solely with line-level supervision, eliminating the need for costly character-level annotations. Key technical contributions include: (1) a Transformer-based multi-instance detector for simultaneous character localization and classification; (2) synthetic data pretraining; (3) dynamic masked consistency learning to enhance robustness; and (4) cross-script transfer strategies. DTLR achieves new state-of-the-art results on CASIA v2 (Chinese), Borg, and Copiale (cipher scripts), demonstrating significantly improved generalization across multilingual and low-resource scripts while drastically reducing annotation dependency.

Enabling character localization across diverse scripts with synthetic pre-training.Improving state-of-the-art performance in Chinese and cipher script recognition.Overcoming challenges in handwritten text recognition (HTR) using detection-based methods.

Latest Papers

What's happening recently
View more

This study addresses the challenging task of character detection and recognition in handwritten forms, which remains difficult due to structural complexity and heavy reliance on extensive manual annotations. The authors propose an end-to-end deep neural network approach that unifies character detection and classification into a single task, thereby eliminating the dependency on manually annotated data inherent in conventional two-stage pipelines. By synthesizing training data using the EMNIST dataset combined with realistic form layouts, the model achieves strong generalization without requiring real-world labeled examples. Evaluated on actual handwritten examination forms, the method attains an overall character recognition accuracy of 88.28%, substantially outperforming existing two-stage approaches.

Deep Neural NetworksEMNISTForm Processing

This study addresses the systemic exclusion of community-produced historical documents—such as Black historical newspapers—from current OCR and document understanding evaluations, which predominantly focus on modern, Western, and institutional materials. By introducing a structural inequality lens into OCR assessment, the research employs a PRISMA-guided systematic review of literature and benchmark datasets from 2006 to 2025, complemented by analyses using vision Transformers, multimodal OCR metrics, and archival empirical data. Findings reveal a critical representational gap: existing evaluations rarely include such marginalized documents and overrely on character-level accuracy, failing to capture layout collapse, font misrecognition, and textual hallucination. This oversight perpetuates the “structural invisibility” and representational harm of historically underrepresented communities, exposing deep-seated institutional biases within evaluation frameworks.

benchmark biasBlack newspapershistorical documents

This study addresses the challenge of word segmentation in Bengali handwritten text captured via mobile devices, where performance is often degraded by paper-ink color variations, complex illumination, and shadow interference. To overcome these limitations, this work proposes a robust automatic segmentation method that transcends conventional color-dependent constraints. A dedicated dataset encompassing multi-colored backgrounds and complex lighting conditions is constructed, and an integrated pipeline combining adaptive thresholding, morphological dilation filtering, and bounding box detection is developed to achieve resilient segmentation. Experimental evaluations on 7,374 test words demonstrate that the proposed approach attains a recall of 90.60%, a precision of 91.80%, and an F1-score of 91.20%. These results indicate a substantial improvement in word segmentation performance for mobile optical character recognition systems operating under unconstrained imaging conditions.

Color IndependentHandwritten BanglaOCR

This work addresses the challenge of accurately segmenting handwritten and printed text in document digitization under the computational constraints of edge devices, where existing deep learning approaches incur high computational costs and are difficult to deploy in lightweight settings. To overcome this, the authors propose a novel lightweight segmentation framework that operates without deep neural networks. The method first extracts semantically coherent text regions via sentence-level connected component segmentation, then introduces a region-aware handwriting descriptor (RHD) to effectively capture the variability inherent in handwriting. Finally, a conventional classifier is employed for efficient discrimination between handwritten and printed text. Evaluated on both the newly curated MAD-HPTS dataset and the public PHD-AS benchmark, the approach outperforms state-of-the-art methods, achieving over 8× inference speedup with only a 1.4% drop in accuracy, thereby substantially reducing computational overhead and enabling practical edge deployment.

document digitizationedge deviceshandwritten text

This study addresses the low OCR accuracy for Old Church Slavonic manuscripts and the reliance of textual phylogeny reconstruction on transcribed texts by proposing a purely vision-driven, end-to-end approach. Through systematic evaluation of conventional OCR systems, machine learning models, and large language models—including GPT-5 and Gemini3-flash—the method integrates an agent-based architecture with retrieval-augmented generation (RAG) post-processing to substantially improve recognition performance. Notably, it achieves the first fully automated, image-only phylogeny reconstruction pipeline, encompassing glyph extraction, clustering, and distance matrix computation. Experimental results demonstrate a character error rate as low as 2–3% and validate the feasibility and effectiveness of the proposed framework on two medieval manuscript corpora.

Church SlavonicManuscript AnalysisOCR

Hot Scholars

WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
FD

Franck Dernoncourt

NLP/ML Researcher. MIT PhD.
Machine LearningNeural NetworksNatural Language Processing
NT

Nan Tang

National Institute of Biological Sciences, Beijing
stem cell biologyaginglung diseases
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
PS

Philip S. Yu

Professor of Computer Science, University of Illinons at Chicago
Data miningDatabasePrivacy