document information extraction

Design and build systems that extract and normalize structured information from documents by parsing full-page text and visual layout to detect and reconstruct table structures (cells, rows, columns and their relationships) and recover entities, clauses, and other fields. This work covers rule-based and learning-based document parsing and comprehension, integration with OCR and multi-page context, table-level extraction and structure recognition, and mapping extracted values to target schemas.

documentinformationextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$203K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

Jul 13, 2025
YD
Yihao Ding
🏛️ The University of Western Australia | The University of Melbourne | Cornell University

This paper presents a systematic survey of recent advances in multimodal large language models (MLLMs) for visually rich document understanding (VRDU). Addressing core challenges—including inadequate integration of textual, visual, and layout modalities, strong reliance on OCR outputs, and limited generalization and robustness—the work comprehensively analyzes MLLM architectures along three dimensions: (1) encoding-fusion mechanisms, (2) module-wise trainability, and (3) staged learning strategies. It comparatively examines OCR-dependent versus OCR-free paradigms and unifies pretraining, instruction tuning, and supervised fine-tuning techniques to propose a novel framework for joint multimodal feature modeling. The survey synthesizes benchmark datasets and representative models, identifies critical bottlenecks—such as layout-aware reasoning gaps and cross-domain instability—and outlines future directions toward efficient, generalizable, and robust VRDU systems. This work serves as a foundational theoretical reference and practical technical guide for the VRDU research community.

Analyzing challenges in OCR-dependent and OCR-free document processingExploring future directions for efficient and robust VRDU systemsSurveying MLLM-based methods for visually rich document understanding

Must-Read Papers

Most classic and influential ideas
View more

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Oct 28, 2024
QZ
Qintong Zhang
🏛️ Shanghai Artificial Intelligence Laboratory | Peking University

This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.

Address challenges in layout detection and multi-modal data integrationConvert unstructured documents into structured machine-readable dataImprove parsing accuracy for complex layouts and high-density text

MonkeyOCR v1.5 Technical Report: Unlocking Robust Document Parsing for Complex Patterns

Nov 13, 2025
JZ
Jiarui Zhang
🏛️ KingSoft Office Zhuiguang AI Lab | Huazhong University of Science and Technology

Real-world documents frequently contain multi-level tables, embedded images/formulas, and cross-page structures, posing significant challenges for existing OCR systems due to insufficient robustness. To address this, we propose a unified vision-language framework featuring a novel two-stage parsing pipeline: (1) a first stage employing image disentanglement and type-guided merging to improve structural fidelity in complex table reconstruction; (2) a second stage leveraging a large multimodal model to jointly predict document layout and reading order, enhanced by a render-and-compare alignment strategy for precise region-wise text, formula, and table recognition. We further introduce a vision-consistency reinforcement learning metric to evaluate and refine recognition quality, substantially improving cross-page and multimodal content handling. Our method achieves state-of-the-art performance on OmniDocBench v1.5, significantly outperforming PPOCR-VL and MinerU 2.5—particularly excelling on visually complex documents.

Addressing complex document layouts with multi-level tables and embedded elementsEnhancing table parsing with embedded images and multi-column reconstructionImproving OCR accuracy for cross-page structures and visual consistency

Low accuracy and structural degradation in table information extraction from scientific literature hinder reliable downstream analysis. Method: We propose a multimodal joint reasoning framework that preserves original table structure by integrating OCR, pre-trained document visual question answering (DocVQA) models, and end-to-end table detection and structure recognition. Our approach jointly models textual semantics and visual layout, explicitly incorporating structured constraints—including row-column relationships, cross-cell semantic dependencies, and mathematical notation—into the QA reasoning process. Contribution/Results: Evaluated on systematic literature review tasks, our method achieves a +12.3% absolute gain in table-related QA accuracy over strong baselines. It demonstrates superior robustness on complex nested tables and multimodal heterogeneous content (e.g., text–formula–figure mixtures), establishing a high-fidelity foundation for structured data extraction in automated literature review systems.

Enhancing extractive question answering accuracyExtracting information from scientific literature tablesPreserving table structure for reliable data extraction

Existing document analysis research predominantly addresses superficial table-centric tasks—such as table detection and layout parsing—while neglecting deep semantic modeling of table-context relationships, thereby hindering cross-paragraph reasoning and consistency analysis. To address this gap, we propose DOTABLER, the first end-to-end framework for joint semantic structure parsing of tables and their surrounding textual context. DOTABLER introduces a unified parsing pipeline integrating layout-aware representation learning, fine-grained semantic matching, and explicit context-table relational modeling. Leveraging a newly constructed PDF-based dataset and domain-adaptive fine-tuning, it enables table-centered document structure modeling and domain-specific retrieval. Evaluated on nearly 4,000 pages of real-world documents, DOTABLER achieves over 90% Precision and F1-score—substantially outperforming strong baselines including GPT-4o—and establishes new state-of-the-art performance in table-context semantic analysis.

Advanced cross-paragraph data interpretation and context-consistent analysisDeep semantic parsing of tables and their contextual associationsPrecise extraction of semantically relevant tables from documents

Existing document layout parsing methods rely on multi-stage pipelines, which suffer from error propagation and lack task-coordinated optimization. To address these limitations, we propose UniDoc—the first end-to-end multilingual vision-language model—unifying layout detection, text recognition, and relational understanding within a single architecture. We further design a scalable synthetic data engine to generate large-scale, high-quality training data spanning 126 languages. Through joint multi-task training and cross-modal alignment, UniDoc significantly enhances robustness and generalization on complex documents. It achieves state-of-the-art performance on OmniDocBench and outperforms the strongest baseline by 7.4 percentage points on our newly constructed multilingual benchmark, XDocParse, demonstrating superior multilingual comprehension and parsing capability.

Enables robust multilingual document parsing across diverse layouts and domainsOvercomes fragmented multi-stage pipelines that cause error propagationUnifies layout detection, text recognition, and relational understanding in a single model

Latest Papers

What's happening recently
View more

Existing document parsing benchmarks are largely confined to single-page or single-task settings, making them inadequate for evaluating semantic continuity, hierarchical structure, and visual fidelity in multi-page documents. This work introduces a realistic, multi-page document parsing benchmark comprising 15 document categories in Chinese and English, with 433 meticulously annotated samples totaling 3,246 pages, enabling end-to-end document-level evaluation for the first time. The benchmark features a fine-grained evaluation protocol that encompasses text, table, and formula recognition; reading order inference; cross-page content consolidation; and heading hierarchy reconstruction. Experimental results reveal that while current models perform reasonably well on basic text extraction, they remain substantially deficient in semantic integration, visual parsing, and structural recovery, thereby establishing a unified and comprehensive foundation for advancing multi-page document understanding.

document parsing benchmarkhierarchical structure recoverymulti-page document parsing

This work addresses severe parsing errors in dense document pages caused by unstable layout assumptions, which lead to mismatches between detector outputs and the input sequence expected by the parser. To resolve this, the authors introduce a lightweight structural refinement module between a DETR-style detector and the parser, performing set-level reasoning using query features, semantic cues, bounding box geometry, and visual evidence. This module jointly decides which instances to retain, refines bounding boxes, and predicts the correct parsing order. By integrating retention-oriented supervision with a difficulty-aware ordering objective, the method significantly enhances layout–parsing interface consistency under complex layouts, reducing the reading order edit distance to 0.024 on OmniDocBench while consistently improving page-level layout quality.

document parsinginstance retentionlayout analysis

Existing end-to-end document parsing methods often suffer from repetitions, hallucinations, and structural inconsistencies in real-world casually captured scenes, primarily due to the scarcity of high-quality full-page supervision data and the absence of structure-aware training mechanisms. This work proposes a co-optimized framework that synergistically enhances data synthesis and model training: it constructs large-scale, structurally diverse full-page supervision data through realistic synthesis based on layout templates and document element composition, and introduces progressive structure-aware training alongside structural token optimization within a billion-parameter multimodal large language model. The proposed approach achieves, for the first time, highly robust end-to-end parsing of real-world captured documents, significantly outperforming existing methods across scanned, digitally generated, and in-the-wild capture scenarios. The project also releases the model, data synthesis pipeline, and a new evaluation benchmark, Wild-OmniDocBench.

data scarcitydocument parsingend-to-end parsing

Existing document parsing and OCR benchmarks struggle to evaluate models’ true capabilities on expert-level complex documents—such as chemical formulas, musical scores, and cross-page tables. To address this gap, this work introduces Dr. DocBench, the first domain-expert-oriented, difficulty-aware document parsing benchmark. Constructed from multilingual book corpora, it employs a parser-failure-driven sampling strategy to curate 4,514 challenging pages, annotated with 65k fine-grained labels covering layout, reading order, hierarchical structure, and domain-specific content across 52 disciplines. Experiments reveal substantial performance degradation among state-of-the-art document parsing systems and general-purpose vision-language models on this benchmark, highlighting their limitations in professional content understanding, modeling of intricate structures, and cross-page contextual reasoning, thereby validating Dr. DocBench’s challenge and efficacy.

benchmarkcomplex layoutsdocument parsing

Hot Scholars

CZ

Chengzhi Zhang

Nanjing University of Science and Technology
Text MiningNatural Language ProcessingScience of Science
PT

Philip Torr

Professor, University of Oxford
Department of Engineering
JY

Jialin Yu

University of Oxford
machine learningartificial intelligencecausalityBayesian inference
SP

Shidong Pan

Postdoctoral Researcher of New York University & Columbia University
Usable Privacy and SecurityPrivacy PolicyResponsible AISoftware Engineering