Score
Designs and implements systems to parse PDF documents into structured, machine-readable content by extracting text, layout, metadata, and embedded elements. Builds and evaluates figure-extraction and segmentation modules that isolate figure images, associate them with captions, and preprocess and clean extracted content for downstream analysis.
Existing PDF parsing tools are primarily designed for academic papers and struggle to accurately process pedagogical documents—such as legal textbooks—that exhibit complex, implicitly structured hierarchies. To address this, we propose a hierarchical text segmentation framework integrating structure-aware preprocessing with large language models (LLMs). Our method jointly leverages OCR-based heading detection, XML structural feature extraction, and contextual semantic modeling to infer implicit heading hierarchies without requiring explicit table-of-contents input. Compared to pure LLM–based or traditional rule-based approaches, our framework significantly reduces false positives and improves segmentation accuracy. When high-quality metadata is available, a supplementary table-of-contents–driven strategy further enhances performance. The source code and benchmark dataset are publicly released to support reproducible research.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
Scientific PDF information extraction tools suffer from inconsistent input formats, opaque “black-box” behavior, poor fault tolerance, and limited format support—hindering literature analysis efficiency for non-NLP researchers. To address these challenges, we propose the first modular information extraction experimental framework specifically designed for scientific PDFs, enabling model-level decoupling, fine-grained intermediate-state visualization, and unified cross-model evaluation. The framework integrates Hugging Face token classifiers, diverse large language models (LLMs), and domain-specific models within a PDF processing pipeline comprising layout-aware parsing, text reconstruction, and semantic alignment. Evaluated on materials science literature review tasks, it significantly reduces model trial-and-error overhead, improves error attribution accuracy, and enhances system interpretability. The platform delivers a debuggable, reusable, plug-and-play rapid prototyping capability for scientific information extraction, empowering domain experts without NLP expertise.
Existing PDF parsing tools lack systematic empirical evaluation on non-academic documents (e.g., patents, financial reports, legal contracts). This paper presents the first comprehensive benchmarking study across six real-world document categories using the DocLayNet dataset, evaluating ten state-of-the-art tools—including PyMuPDF, Nougat, and TATR—on text extraction and table detection. Results reveal strong document-type dependency in tool performance: learning-based models (Nougat, TATR) consistently outperform rule-based approaches on complex layouts; PyMuPDF achieves the best overall text extraction accuracy; TATR leads in table detection across four document types; Camelot exhibits superior adaptability to tender documents; and Nougat excels on scientific and patent documents. By establishing a rigorous, domain-diverse evaluation framework, this work fills a critical gap in empirical assessment of PDF parsing for non-academic domains and provides evidence-based guidance for context-aware tool selection.
Existing document parsing methods suffer from narrow benchmark coverage and oversimplified evaluation protocols, resulting in unrealistic and non-comprehensive assessments. Method: We introduce the first comprehensive, multi-source PDF benchmark—encompassing nine real-world document types (e.g., academic papers, textbooks, slides)—and propose a unified, fine-grained evaluation framework with 19 layout categories and 14 attribute classes. Leveraging a high-quality, human-annotated dataset, we systematically compare modular pipeline approaches against multimodal end-to-end models. Contribution/Results: Our empirical analysis exposes critical limitations in current methods’ ability to handle document diversity and structural generalization. To foster reproducibility and community advancement, we publicly release the benchmark dataset, source code, and evaluation toolkit—establishing a new standard for rigorous, cross-model, cross-module, and cross-document-type evaluation in document parsing.
Existing PDF parsers often fail to capture critical visual content, erroneously extract irrelevant images, and struggle to accurately associate figures with their captions—limitations that significantly hinder the performance of multimodal retrieval-augmented generation (RAG) systems. To address these challenges, this work proposes a lightweight, production-oriented visual PDF parsing framework that, for the first time in a production setting, integrates spatial heuristics, document layout analysis, and semantic similarity to achieve robust, low-latency detection of visual elements and precise figure–caption alignment. Evaluated on both public and internal datasets, the method attains detection accuracy of at least 96% and caption association accuracy of 93%. When deployed as a RAG preprocessing module, it substantially outperforms current state-of-the-art approaches while reducing inference latency by more than twofold.
Existing benchmarks struggle to systematically evaluate the impact of multimodal inputs on structured information extraction. This work proposes the first multimodal extraction benchmark tailored for government forms, generating diverse documents with deterministic ground truth through procedural PDF templates and a reverse annotation pipeline. Each document is provided in four standardized input formats—plain text, layout-preserving text, image, and multimodal—to enable rigorous ablation studies. A three-stage quality control mechanism and compliance-based evaluation of structured outputs ensure high data fidelity and reliable assessment. Experiments reveal that small models (<4B parameters) are primarily limited by structural adherence; fine-tuning a 2B model yields an 81-percentage-point improvement; layout-preserving text consistently outperforms image inputs by 3–18 percentage points; and the benchmark exhibits maximal discriminative power within the 60–95% accuracy range.
This study addresses the challenges posed by heterogeneous content—comprising text, tables, and images—in financial PDF documents for retrieval-augmented generation (RAG) systems, an area lacking systematic evaluation of parsing and chunking strategies. The work presents the first end-to-end empirical investigation of PDF processing pipelines in the context of financial document question answering. It introduces a new benchmark, TableQuest, integrates existing datasets, and systematically evaluates diverse PDF parsers and text chunking methods—including overlap mechanisms—on downstream QA performance. The findings reveal a strong correlation between preservation of document structure and answer accuracy. To support reproducibility and future research, the authors release their benchmark publicly and provide actionable, empirically grounded guidelines for designing effective RAG pipelines tailored to complex financial documents.
Existing document parsing benchmarks are largely confined to single-page or single-task settings, making them inadequate for evaluating semantic continuity, hierarchical structure, and visual fidelity in multi-page documents. This work introduces a realistic, multi-page document parsing benchmark comprising 15 document categories in Chinese and English, with 433 meticulously annotated samples totaling 3,246 pages, enabling end-to-end document-level evaluation for the first time. The benchmark features a fine-grained evaluation protocol that encompasses text, table, and formula recognition; reading order inference; cross-page content consolidation; and heading hierarchy reconstruction. Experimental results reveal that while current models perform reasonably well on basic text extraction, they remain substantially deficient in semantic integration, visual parsing, and structural recovery, thereby establishing a unified and comprehensive foundation for advancing multi-page document understanding.