character segmentation

Designs and implements algorithms or pipelines that detect and extract individual character regions from input images (e.g., scanned documents, scene text, or handwriting), separate touching or occluded characters, normalize character scale and orientation, and produce character segments prepared for downstream recognition.

charactersegmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Low-quality invoice images—characterized by complex table structures, severe noise, and heterogeneous layouts—significantly degrade OCR accuracy. To address this, we propose an end-to-end OCR-driven pipeline for tabular data extraction. Our method introduces a dynamic image preprocessing mechanism to enhance readability of degraded invoices; designs an adaptive table boundary detection and row-column mapping algorithm to robustly localize non-standard tables and semantically align cells; and integrates Tesseract OCR with customized post-processing logic for accurate text recognition and structured reconstruction. Experiments on a real-world invoice dataset demonstrate substantial improvements: +12.7% in field-level accuracy and enhanced layout consistency. The pipeline enables high-precision financial automation and digital archival, exhibiting strong engineering deployability in production environments.

Automating financial workflows via invoice digitizationExtracting structured tabular data from noisy invoicesImproving accuracy in OCR-based table recognition

Digitization of Document and Information Extraction using OCR

Jun 11, 2025
RS
Rasha Sinha
🏛️ RV College of Engineering

To address the limited flexibility and weak semantic understanding in information extraction from mixed scanned images and natively digital documents, this paper proposes a two-stage OCR-LLM collaborative framework. In the first stage, multi-engine OCR combined with layout-aware parsing (e.g., PDFMiner/LayoutParser) performs preliminary text and structural extraction. In the second stage, a fine-tuned large language model (LLM), enhanced by context-aware prompt engineering, achieves cross-format, multi-layout semantic parsing of key fields and generates confidence scores. The work introduces the novel “OCR pre-extraction + LLM post-parsing” paradigm, integrating visual layout cues with linguistic context to resolve format-induced ambiguities. Experiments demonstrate a 32% improvement in key-field extraction accuracy, support for over ten document types, a 27% reduction in average latency, and significantly enhanced layout robustness and generalization capability.

Combining OCR and LLMs for structured, context-aware data extractionEvaluating OCR tools for accuracy, layout recognition, and speedExtracting accurate text from mixed scanned and digital documents

Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR

Aug 29, 2025
SV
Shashank Vempati
🏛️ Typeface | University of Maryland, College Park | Tata 1mg | Vellore Institute of Technology | Indian Institute of Technology Delhi

To address the high character/word segmentation errors and insufficient contextual modeling in traditional OCR, this paper proposes a line-level OCR paradigm that bypasses explicit character and word segmentation and performs end-to-end recognition directly on full text lines. Methodologically, we introduce a unified sequence-to-sequence framework integrating object detection with deep language modeling. We provide the first systematic empirical validation of the advantages of line-level modeling and release LineOCR, the first fine-grained annotation dataset specifically designed for line-level training and evaluation (251 pages of English documents). Experiments demonstrate that our approach achieves a 5.4% absolute improvement in end-to-end accuracy and a 4× speedup in inference latency, substantially alleviating bottlenecks inherent in conventional “segment-then-recognize” pipelines. This work advances OCR toward a unified perception-and-understanding paradigm.

Addresses word segmentation errors in OCR systemsEnhances language model context and improves accuracyProposes line-level OCR to bypass word detection issues

To address the low OCR accuracy, poor layout adaptability, and inefficiency in large-scale document processing inherent in traditional RPA systems for immigration document handling, this paper proposes an LLM-augmented intelligent RPA framework. The method pioneers the integration of fine-tuned large language models (LLMs) across the entire RPA pipeline, synergistically combining OCR engines with rule-enhanced workflows and context-aware textual post-processing to enable fuzzy character correction, complex layout parsing, and end-to-end structured information extraction. Experiments demonstrate that ID data extraction time is reduced to an average of 9.94 seconds—up to 94% faster than UiPath and Automation Anywhere—while accuracy and cross-document robustness are significantly improved. This work establishes a scalable technical paradigm for automated understanding of high-noise, multi-layout government documents.

Automation Tool EffectivenessDocument Processing EfficiencyImmigration File Information Extraction

PreP-OCR: A Complete Pipeline for Document Image Restoration and Enhanced OCR Accuracy

May 26, 2025
SG
Shuhao Guan
🏛️ University College Dublin | Trinity College Dublin | University of Toronto | Shanghai University

To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.

Combining image restoration with OCR error correctionImproving text extraction from degraded historical documentsReducing character error rates in multilingual historical texts

Latest Papers

What's happening recently
View more

This work addresses the challenge of accurately segmenting handwritten and printed text in document digitization under the computational constraints of edge devices, where existing deep learning approaches incur high computational costs and are difficult to deploy in lightweight settings. To overcome this, the authors propose a novel lightweight segmentation framework that operates without deep neural networks. The method first extracts semantically coherent text regions via sentence-level connected component segmentation, then introduces a region-aware handwriting descriptor (RHD) to effectively capture the variability inherent in handwriting. Finally, a conventional classifier is employed for efficient discrimination between handwritten and printed text. Evaluated on both the newly curated MAD-HPTS dataset and the public PHD-AS benchmark, the approach outperforms state-of-the-art methods, achieving over 8× inference speedup with only a 1.4% drop in accuracy, thereby substantially reducing computational overhead and enabling practical edge deployment.

document digitizationedge deviceshandwritten text

This study addresses the challenging task of character detection and recognition in handwritten forms, which remains difficult due to structural complexity and heavy reliance on extensive manual annotations. The authors propose an end-to-end deep neural network approach that unifies character detection and classification into a single task, thereby eliminating the dependency on manually annotated data inherent in conventional two-stage pipelines. By synthesizing training data using the EMNIST dataset combined with realistic form layouts, the model achieves strong generalization without requiring real-world labeled examples. Evaluated on actual handwritten examination forms, the method attains an overall character recognition accuracy of 88.28%, substantially outperforming existing two-stage approaches.

Deep Neural NetworksEMNISTForm Processing

This study addresses the inefficiency of manual verification processes for handwritten signatures in Swiss popular initiatives, highlighting the urgent need for automation. The authors propose the first end-to-end analysis pipeline that integrates template-based line segmentation, OCR-based text recognition, and vision-based handwriting retrieval, pioneering the application of handwriting retrieval to detect potentially duplicated submissions within signature lists. Experimental results demonstrate that generic OCR systems exhibit limited performance on short handwritten names, achieving a character error rate (CER) of 29.6%, whereas the handwriting retrieval approach shows greater promise for deduplication, attaining a mean average precision (mAP) of 50.6%. These findings validate the method’s effectiveness and novelty in supporting human reviewers during the verification process.

handwriting extractionOCRpopular initiatives

Hot Scholars

IK

Injung Kim

Professor, Handong Global University
AIdeep learningimage analysis and synthesisspeech synthesis
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
SS

Shuzheng Si

Tsinghua University
Natural Language ProcessingLarge Language Models
MS

Maosong Sun

Professor of Computer Science and Technology, Tsinghua University
Natural Language ProcessingArtificial IntelligenceSocial Computing
PG

Pan Gao

Professor, Nanjing University of Aeronautics and Astronautics;
Image/Video/Point cloudsdeep learningMultimedia