Score
Design and build systems that extract and normalize structured information from documents by parsing full-page text and visual layout to detect and reconstruct table structures (cells, rows, columns and their relationships) and recover entities, clauses, and other fields. This work covers rule-based and learning-based document parsing and comprehension, integration with OCR and multi-page context, table-level extraction and structure recognition, and mapping extracted values to target schemas.
Text-to-structured generation (e.g., tables, knowledge graphs, charts) for agent-centric AI is a foundational infrastructure enabling context-aware retrieval and autonomous reasoning, yet suffers from fragmented methodologies, scarce standardized datasets, and inconsistent evaluation protocols. Method: We conduct a systematic literature review integrating techniques from NLP, information extraction, knowledge representation, and machine learning to establish the first holistic analytical framework—comprising task taxonomy, benchmark dataset inventory, and unified evaluation metrics. Contribution/Results: We introduce the first general-purpose evaluation framework for structured output generation, explicitly identifying methodological limitations and core challenges (e.g., fidelity, composability, and reasoning-aware assessment). We comprehensively map research gaps and affirm the centrality of this direction in next-generation AI systems, providing both theoretical grounding and practical guidance for future algorithmic development and empirical validation.
This paper presents a systematic survey of recent advances in multimodal large language models (MLLMs) for visually rich document understanding (VRDU). Addressing core challenges—including inadequate integration of textual, visual, and layout modalities, strong reliance on OCR outputs, and limited generalization and robustness—the work comprehensively analyzes MLLM architectures along three dimensions: (1) encoding-fusion mechanisms, (2) module-wise trainability, and (3) staged learning strategies. It comparatively examines OCR-dependent versus OCR-free paradigms and unifies pretraining, instruction tuning, and supervised fine-tuning techniques to propose a novel framework for joint multimodal feature modeling. The survey synthesizes benchmark datasets and representative models, identifies critical bottlenecks—such as layout-aware reasoning gaps and cross-domain instability—and outlines future directions toward efficient, generalizable, and robust VRDU systems. This work serves as a foundational theoretical reference and practical technical guide for the VRDU research community.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
Real-world documents frequently contain multi-level tables, embedded images/formulas, and cross-page structures, posing significant challenges for existing OCR systems due to insufficient robustness. To address this, we propose a unified vision-language framework featuring a novel two-stage parsing pipeline: (1) a first stage employing image disentanglement and type-guided merging to improve structural fidelity in complex table reconstruction; (2) a second stage leveraging a large multimodal model to jointly predict document layout and reading order, enhanced by a render-and-compare alignment strategy for precise region-wise text, formula, and table recognition. We further introduce a vision-consistency reinforcement learning metric to evaluate and refine recognition quality, substantially improving cross-page and multimodal content handling. Our method achieves state-of-the-art performance on OmniDocBench v1.5, significantly outperforming PPOCR-VL and MinerU 2.5—particularly excelling on visually complex documents.
Low accuracy and structural degradation in table information extraction from scientific literature hinder reliable downstream analysis. Method: We propose a multimodal joint reasoning framework that preserves original table structure by integrating OCR, pre-trained document visual question answering (DocVQA) models, and end-to-end table detection and structure recognition. Our approach jointly models textual semantics and visual layout, explicitly incorporating structured constraints—including row-column relationships, cross-cell semantic dependencies, and mathematical notation—into the QA reasoning process. Contribution/Results: Evaluated on systematic literature review tasks, our method achieves a +12.3% absolute gain in table-related QA accuracy over strong baselines. It demonstrates superior robustness on complex nested tables and multimodal heterogeneous content (e.g., text–formula–figure mixtures), establishing a high-fidelity foundation for structured data extraction in automated literature review systems.
Existing document analysis research predominantly addresses superficial table-centric tasks—such as table detection and layout parsing—while neglecting deep semantic modeling of table-context relationships, thereby hindering cross-paragraph reasoning and consistency analysis. To address this gap, we propose DOTABLER, the first end-to-end framework for joint semantic structure parsing of tables and their surrounding textual context. DOTABLER introduces a unified parsing pipeline integrating layout-aware representation learning, fine-grained semantic matching, and explicit context-table relational modeling. Leveraging a newly constructed PDF-based dataset and domain-adaptive fine-tuning, it enables table-centered document structure modeling and domain-specific retrieval. Evaluated on nearly 4,000 pages of real-world documents, DOTABLER achieves over 90% Precision and F1-score—substantially outperforming strong baselines including GPT-4o—and establishes new state-of-the-art performance in table-context semantic analysis.
Existing document layout parsing methods rely on multi-stage pipelines, which suffer from error propagation and lack task-coordinated optimization. To address these limitations, we propose UniDoc—the first end-to-end multilingual vision-language model—unifying layout detection, text recognition, and relational understanding within a single architecture. We further design a scalable synthetic data engine to generate large-scale, high-quality training data spanning 126 languages. Through joint multi-task training and cross-modal alignment, UniDoc significantly enhances robustness and generalization on complex documents. It achieves state-of-the-art performance on OmniDocBench and outperforms the strongest baseline by 7.4 percentage points on our newly constructed multilingual benchmark, XDocParse, demonstrating superior multilingual comprehension and parsing capability.
Existing document parsing benchmarks are largely confined to single-page or single-task settings, making them inadequate for evaluating semantic continuity, hierarchical structure, and visual fidelity in multi-page documents. This work introduces a realistic, multi-page document parsing benchmark comprising 15 document categories in Chinese and English, with 433 meticulously annotated samples totaling 3,246 pages, enabling end-to-end document-level evaluation for the first time. The benchmark features a fine-grained evaluation protocol that encompasses text, table, and formula recognition; reading order inference; cross-page content consolidation; and heading hierarchy reconstruction. Experimental results reveal that while current models perform reasonably well on basic text extraction, they remain substantially deficient in semantic integration, visual parsing, and structural recovery, thereby establishing a unified and comprehensive foundation for advancing multi-page document understanding.
This work addresses severe parsing errors in dense document pages caused by unstable layout assumptions, which lead to mismatches between detector outputs and the input sequence expected by the parser. To resolve this, the authors introduce a lightweight structural refinement module between a DETR-style detector and the parser, performing set-level reasoning using query features, semantic cues, bounding box geometry, and visual evidence. This module jointly decides which instances to retain, refines bounding boxes, and predicts the correct parsing order. By integrating retention-oriented supervision with a difficulty-aware ordering objective, the method significantly enhances layout–parsing interface consistency under complex layouts, reducing the reading order edit distance to 0.024 on OmniDocBench while consistently improving page-level layout quality.
Existing end-to-end document parsing methods often suffer from repetitions, hallucinations, and structural inconsistencies in real-world casually captured scenes, primarily due to the scarcity of high-quality full-page supervision data and the absence of structure-aware training mechanisms. This work proposes a co-optimized framework that synergistically enhances data synthesis and model training: it constructs large-scale, structurally diverse full-page supervision data through realistic synthesis based on layout templates and document element composition, and introduces progressive structure-aware training alongside structural token optimization within a billion-parameter multimodal large language model. The proposed approach achieves, for the first time, highly robust end-to-end parsing of real-world captured documents, significantly outperforming existing methods across scanned, digitally generated, and in-the-wild capture scenarios. The project also releases the model, data synthesis pipeline, and a new evaluation benchmark, Wild-OmniDocBench.
Existing document parsing and OCR benchmarks struggle to evaluate models’ true capabilities on expert-level complex documents—such as chemical formulas, musical scores, and cross-page tables. To address this gap, this work introduces Dr. DocBench, the first domain-expert-oriented, difficulty-aware document parsing benchmark. Constructed from multilingual book corpora, it employs a parser-failure-driven sampling strategy to curate 4,514 challenging pages, annotated with 65k fine-grained labels covering layout, reading order, hierarchical structure, and domain-specific content across 52 disciplines. Experiments reveal substantial performance degradation among state-of-the-art document parsing systems and general-purpose vision-language models on this benchmark, highlighting their limitations in professional content understanding, modeling of intricate structures, and cross-page contextual reasoning, thereby validating Dr. DocBench’s challenge and efficacy.