document parsing

Designs and implements systems and pipelines that ingest and preprocess semi-structured and unstructured documents (e.g., PDF, HTML, webpages, plain text), analyze layout and structure, segment or chunk content, and extract and normalize textual and layout features into structured outputs (tables, JSON, labeled spans) for downstream use. Builds and evaluates parsing algorithms, layout-analysis and document-processing components, content/web parsing and matching modules, and end-to-end document understanding workflows that handle encoding/OCR variability and preserve semantic regions.

documentparsing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.5
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction

Oct 28, 2024
QZ
Qintong Zhang
🏛️ Shanghai Artificial Intelligence Laboratory | Peking University

This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.

Address challenges in layout detection and multi-modal data integrationConvert unstructured documents into structured machine-readable dataImprove parsing accuracy for complex layouts and high-density text

HiPS: Hierarchical PDF Segmentation of Textbooks

Aug 31, 2025
SW
Sabine Wehnert
🏛️ Otto von Guericke University | Leibniz Institute for Educational Media | Georg Eckert Institute

Existing PDF parsing tools are primarily designed for academic papers and struggle to accurately process pedagogical documents—such as legal textbooks—that exhibit complex, implicitly structured hierarchies. To address this, we propose a hierarchical text segmentation framework integrating structure-aware preprocessing with large language models (LLMs). Our method jointly leverages OCR-based heading detection, XML structural feature extraction, and contextual semantic modeling to infer implicit heading hierarchies without requiring explicit table-of-contents input. Compared to pure LLM–based or traditional rule-based approaches, our framework significantly reduces false positives and improves segmentation accuracy. When high-quality metadata is available, a supplementary table-of-contents–driven strategy further enhances performance. The source code and benchmark dataset are publicly released to support reproducible research.

Evaluating TOC-based and structure-aware methods for textbook analysisHierarchical segmentation of complex structured PDF documentsImproving parsing accuracy with preprocessing and LLM integration

Low-quality invoice images—characterized by complex table structures, severe noise, and heterogeneous layouts—significantly degrade OCR accuracy. To address this, we propose an end-to-end OCR-driven pipeline for tabular data extraction. Our method introduces a dynamic image preprocessing mechanism to enhance readability of degraded invoices; designs an adaptive table boundary detection and row-column mapping algorithm to robustly localize non-standard tables and semantically align cells; and integrates Tesseract OCR with customized post-processing logic for accurate text recognition and structured reconstruction. Experiments on a real-world invoice dataset demonstrate substantial improvements: +12.7% in field-level accuracy and enhanced layout consistency. The pipeline enables high-precision financial automation and digital archival, exhibiting strong engineering deployability in production environments.

Automating financial workflows via invoice digitizationExtracting structured tabular data from noisy invoicesImproving accuracy in OCR-based table recognition

This work addresses the limitations of existing methods in question answering over semi-structured documents containing tables, figures, and hierarchical paragraphs, which often suffer from fragmented OCR outputs, inadequate hierarchical modeling, and insufficient cross-region alignment. To overcome these challenges, we propose MoDora, a novel system that reconstructs OCR results through local alignment aggregation, explicitly models inter-component hierarchical and spatial relationships via a Component Correlation Tree (CCTree), and introduces a question-type-aware hybrid retrieval mechanism that fuses semantic and layout information. Evaluated on multiple benchmarks, MoDora significantly outperforms current state-of-the-art approaches, achieving absolute accuracy gains of 5.97% to 61.07% and enabling precise, layout- and semantics-driven document understanding.

hierarchical structureinformation retrievallayout-aware representation

The Design of an LLM-powered Unstructured Analytics System

Sep 01, 2024
EA
Eric Anderson
🏛️ Aryn, Inc.

To address the low accuracy and poor interpretability of semantic analysis for unstructured documents, this paper proposes Aryn: a task-oriented LLM system architecture paradigm. Methodologically, Aryn introduces a novel deep coupling between a declarative document processing engine (Sycamore) and a natural-language-to-code query planner (Luna), integrated with the DocParse parser to form an end-to-end analytical pipeline. This design overcomes the accuracy limitations of conventional RAG systems in complex semantic reasoning, enabling natural-language querying, automatic generation of executable semantic plans, and full execution provenance with intermediate-result visualization. Evaluated on real-world NTSB aviation accident reports, Aryn achieves significantly higher accuracy than baseline RAG approaches while substantially improving user trust and debugging capability.

Data ExtractionEfficiency and Accuracy ImprovementNon-structured Information Analysis

Latest Papers

What's happening recently
View more

This work addresses the challenge of accurately and efficiently digitizing complex documents containing handwritten content, irregular tables, and heterogeneous layouts—tasks that remain difficult for conventional OCR systems and current large language models. The authors propose an interactive document digitization system that integrates layout-aware parsing, OCR, and a large language model, enhanced by a user-in-the-loop correction propagation mechanism. Leveraging layout-aware inference, the system automatically generalizes user edits or natural language instructions applied to a local region to structurally similar regions across the document. In a user study (n=12), this approach significantly improved correction efficiency, reduced repetitive manual operations, and enabled more controllable and effective reconstruction of document structure and content.

document digitizationhandwritten contentheterogeneous layouts

This work addresses the limitation of existing document analysis systems that flatten complex documents into plain text, thereby discarding critical hierarchical structures such as sections, tables, and figures, which hinders effective filtering and in-depth analysis. To overcome this, the authors propose a structure-aware document understanding framework that integrates full document hierarchy into semantic indexing and analysis. The approach constructs a hierarchical document tree by parsing the original layout, leverages large language models to generate structure-aware semantic representations, and introduces a multi-view interactive web interface enabling precise natural language–driven retrieval and question answering. Experiments demonstrate significant improvements in retrieval accuracy and question-answering performance on diverse complex documents, including academic papers, technical manuals, and financial reports. The code and a live demo system are publicly released.

complex documentsdocument analysishierarchical structure

This study addresses the lack of efficient, automated methods for structuring heterogeneous real estate questionnaire documents. To this end, the authors propose an end-to-end information extraction framework that first categorizes documents into structural types using K-Means clustering and text classification. Subsequently, it leverages the DeepSeek-R1 large language model enhanced with prompt engineering to accurately extract 35 predefined attributes from complex document formats—including checkboxes and scanned images—and outputs them as structured JSON. Evaluated on a dataset of 2,781 documents, the method produced 2,766 unique property records. Downstream validation demonstrated a Jaccard similarity of 0.82, marking the first high-precision, scalable solution for structured information extraction from such challenging real estate documentation.

automated processingheterogeneous documentsproperty metadata

This work addresses the gap between research and production deployment in large-scale multi-page document processing by proposing a microservice architecture tailored for high-throughput scenarios, integrating a multi-stage pipeline of document classification, optical character recognition (OCR), and large language model (LLM) inference. The system employs a hybrid classification strategy, decouples GPU-based inference from CPU-driven orchestration, leverages asynchronous I/O, and supports independent horizontal scaling, enabling stable processing of thousands of documents per hour. Empirical analysis reveals that OCR constitutes the primary bottleneck in end-to-end latency and that system concurrency is constrained by the inference capacity of shared GPUs rather than the number of nodes. This study offers a reusable, efficient deployment paradigm for industrial-scale document understanding systems.

Document AILLM pipelinesmicroservice architecture

Hot Scholars

VG

Vivek Gupta

Assistant Professor of Computer Science, Arizona State University
Artificial IntelligenceNatural Language ProcessingLarge Language ModelsInformation Retrieval
WZ

Wentao Zhang

Institute of Physics, Chinese Academy of Sciences
photoemissionsuperconductivitycupratehtsc
SH

Shi Han

Microsoft Research Asia
Software AnalyticsMachine LearningData Mining
DZ

Dongmei Zhang

Microsoft Research
Software EngineeringMachine LearningInformation Visualization
MZ

Mengyu Zhou

Microsoft Research
Data analyticsNatural Language ProcessingNetwork ScienceHuman Behaviors