Score
Designs and implements end-to-end OCR processing pipelines that preprocess visual inputs, perform optical character recognition to extract text, and integrate or merge OCR outputs across pages and sources. Builds components to ground extracted text to structured data schemas and to evaluate OCR accuracy and layout-aware extraction quality.
该文通过系统性文献回顾,评估了近十年OCR技术的发展,分析了97篇相关研究,探讨了解决多语言处理和复杂数据格式的方法及挑战。
This paper addresses the insufficient modeling and fusion of textual characteristics in vision-text processing. To tackle this, we propose a systematic solution: (1) We introduce VTPBench, the first end-to-end benchmark covering detection, recognition, reconstruction, and editing tasks; (2) We design VTPScore, an MLLM-based, semantics-aware automatic evaluation metric enabling fair, cross-task quantitative assessment; and (3) Through empirical analysis of over 20 state-of-the-art models, we identify pervasive deficiencies in semantic consistency and geometric fidelity. Our contributions are threefold: (1) The first unified, full-spectrum evaluation benchmark for vision-text processing tasks; (2) The first MLLM-driven evaluation framework explicitly supporting semantic understanding; and (3) Open-sourced, reproducible diagnostic tools and resources that provide both theoretical foundations and practical paradigms for text-specific modeling.
Low-quality invoice images—characterized by complex table structures, severe noise, and heterogeneous layouts—significantly degrade OCR accuracy. To address this, we propose an end-to-end OCR-driven pipeline for tabular data extraction. Our method introduces a dynamic image preprocessing mechanism to enhance readability of degraded invoices; designs an adaptive table boundary detection and row-column mapping algorithm to robustly localize non-standard tables and semantically align cells; and integrates Tesseract OCR with customized post-processing logic for accurate text recognition and structured reconstruction. Experiments on a real-world invoice dataset demonstrate substantial improvements: +12.7% in field-level accuracy and enhanced layout consistency. The pipeline enables high-precision financial automation and digital archival, exhibiting strong engineering deployability in production environments.
To address the limited flexibility and weak semantic understanding in information extraction from mixed scanned images and natively digital documents, this paper proposes a two-stage OCR-LLM collaborative framework. In the first stage, multi-engine OCR combined with layout-aware parsing (e.g., PDFMiner/LayoutParser) performs preliminary text and structural extraction. In the second stage, a fine-tuned large language model (LLM), enhanced by context-aware prompt engineering, achieves cross-format, multi-layout semantic parsing of key fields and generates confidence scores. The work introduces the novel “OCR pre-extraction + LLM post-parsing” paradigm, integrating visual layout cues with linguistic context to resolve format-induced ambiguities. Experiments demonstrate a 32% improvement in key-field extraction accuracy, support for over ten document types, a 27% reduction in average latency, and significantly enhanced layout robustness and generalization capability.
To address low OCR accuracy caused by degradation in historical document images, this paper proposes a two-stage end-to-end optimization framework. In the first stage, a U-Net–based image restoration model is trained on a synthetically generated multi-degradation dataset to jointly optimize visual clarity and linguistic consistency. In the second stage, a semantic-aware ByT5 model performs post-OCR error correction, enhanced by a multi-directional block extraction and fusion mechanism tailored for large-format documents. The key innovations include the first joint optimization of image restoration quality and text semantic consistency, and the construction of the first cross-lingual (English/French/Spanish) synthetic dataset for historical text. Evaluated on 13,831 pages of real historical documents, the framework reduces character error rate by 63.9–70.3% over baseline OCR systems, demonstrating substantial improvement.
This study investigates the impact of optical character recognition (OCR) information injection on the performance of vision-language models (VLMs), focusing on three representative architectures: Qwen3-VL, Phi-4, and InternVL3.5. Through causal intervention, activation difference analysis, principal component analysis (PCA), and cross-dataset directional transfer experiments, the work reveals for the first time that OCR signals propagate via a low-dimensional shared pathway, with the first principal component (PC1) accounting for 72.9% of the variance and demonstrating strong cross-dataset generalization. Notably, ablating the OCR module in Qwen3-VL-4B improves counting task accuracy by up to 6.9 percentage points, suggesting that OCR can interfere with non-text visual reasoning. These findings highlight the dual-edged role of OCR in multimodal fusion—beneficial for text-related tasks yet potentially detrimental to broader visual understanding.
This work addresses the limitations of existing OCR systems, which predominantly focus on text-centric tasks and struggle to effectively process multimodal content in visually dense images such as charts and web pages. To bridge this gap, we propose OCRVerse—the first end-to-end, full-stack OCR framework that unifies modeling for both text- and vision-centric tasks. We construct a large-scale dataset encompassing diverse document types and complex visual layouts, and introduce a two-stage supervised fine-tuning (SFT) followed by reinforcement learning (RL) training strategy, augmented with a domain-adaptive reward mechanism to mitigate inter-domain data conflicts and output format discrepancies. Experimental results demonstrate that OCRVerse achieves performance on par with current state-of-the-art open- and closed-source large models across both task categories.
本文通过微调SmolDocling模型直接从文档图像中端到端提取键值对,解决了传统流程中的多阶段错误传播问题。
This work addresses the gap between research and production deployment in large-scale multi-page document processing by proposing a microservice architecture tailored for high-throughput scenarios, integrating a multi-stage pipeline of document classification, optical character recognition (OCR), and large language model (LLM) inference. The system employs a hybrid classification strategy, decouples GPU-based inference from CPU-driven orchestration, leverages asynchronous I/O, and supports independent horizontal scaling, enabling stable processing of thousands of documents per hour. Empirical analysis reveals that OCR constitutes the primary bottleneck in end-to-end latency and that system concurrency is constrained by the inference capacity of shared GPUs rather than the number of nodes. This study offers a reusable, efficient deployment paradigm for industrial-scale document understanding systems.
研究评估了历史文献数字化流程,采用直接提取、大语言模型后校正及分块提取法,解决了OCR错误传播问题,提高信息检索准确性。
This study addresses the data privacy concerns, high costs, and energy consumption associated with deploying commercial closed-source vision-language models (VLMs) for heritage archive digitization. We propose an autonomous and controllable structured extraction framework based on a lightweight open-source VLM. Methodologically, utilizing a 7B-parameter base model, we design a constraint-aware protocol that integrates classical image preprocessing with multi-stage fine-tuning techniques, while quantifying the independent contribution of each module to document OCR-to-JSON extraction accuracy. Experimental results demonstrate that the adapted lightweight model can efficiently replace manual annotation or black-box systems, achieving compliant and precise structured data extraction. This work presents a novel paradigm that balances performance with controllability for the digitization of sensitive archival materials.
研究针对高风险公共部门应用中的信息提取问题,通过评估开源OCR引擎、大语言模型和视觉-语言模型在处理复杂文档任务上的表现,揭示了现有模型的局限性和影响因素。