Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices

📅 2025-07-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Low-quality invoice images—characterized by complex table structures, severe noise, and heterogeneous layouts—significantly degrade OCR accuracy. To address this, we propose an end-to-end OCR-driven pipeline for tabular data extraction. Our method introduces a dynamic image preprocessing mechanism to enhance readability of degraded invoices; designs an adaptive table boundary detection and row-column mapping algorithm to robustly localize non-standard tables and semantically align cells; and integrates Tesseract OCR with customized post-processing logic for accurate text recognition and structured reconstruction. Experiments on a real-world invoice dataset demonstrate substantial improvements: +12.7% in field-level accuracy and enhanced layout consistency. The pipeline enables high-precision financial automation and digital archival, exhibiting strong engineering deployability in production environments.

Technology Category

Natural Language Processing: Information ExtractionSearch and Optimization: Distributed SearchComputer Vision: Learning & Optimization for CV

Application Category

Economics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphs
📝 Abstract
This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, align, and extract structured tabular data from scanned invoice documents. Our approach includes dynamic preprocessing, table boundary detection, and row-column mapping, optimized for noisy and non-standard invoice formats. The resulting pipeline significantly improves data extraction accuracy and consistency, supporting real-world use cases such as automated financial workflows and digital archiving.
Problem

Research questions and friction points this paper is trying to address.

Extracting structured tabular data from noisy invoices
Improving accuracy in OCR-based table recognition
Automating financial workflows via invoice digitization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses Tesseract OCR for text recognition
Implements custom post-processing for table extraction
Optimizes for noisy, non-standard invoice formats
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
P
Parshva D. Patel