OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

📅 2024-12-10
🏛️ arXiv.org
📈 Citations: 2
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing document parsing methods suffer from narrow benchmark coverage and oversimplified evaluation protocols, resulting in unrealistic and non-comprehensive assessments. Method: We introduce the first comprehensive, multi-source PDF benchmark—encompassing nine real-world document types (e.g., academic papers, textbooks, slides)—and propose a unified, fine-grained evaluation framework with 19 layout categories and 14 attribute classes. Leveraging a high-quality, human-annotated dataset, we systematically compare modular pipeline approaches against multimodal end-to-end models. Contribution/Results: Our empirical analysis exposes critical limitations in current methods’ ability to handle document diversity and structural generalization. To foster reproducibility and community advancement, we publicly release the benchmark dataset, source code, and evaluation toolkit—establishing a new standard for rigorous, cross-model, cross-module, and cross-document-type evaluation in document parsing.

Technology Category

Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsMachine Learning: Large Multimodal Models (LMMs)Computer Vision: Multi-modal Vision

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Document content extraction is crucial in computer vision, especially for meeting the high-quality data needs of large language models (LLMs) and retrieval-augmented generation (RAG) technologies. However, current document parsing methods suffer from significant limitations in terms of diversity and comprehensive evaluation. To address these challenges, we introduce OmniDocBench, a novel multi-source benchmark designed to advance automated document content extraction. OmniDocBench includes a meticulously curated and annotated high-quality evaluation dataset comprising nine diverse document types, such as academic papers, textbooks, slides, among others. Our benchmark provides a flexible and comprehensive evaluation framework with 19 layout category labels and 14 attribute labels, enabling multi-level assessments across entire datasets, individual modules, or specific data types. Using OmniDocBench, we perform an exhaustive comparative analysis of existing modular pipelines and multimodal end-to-end methods, highlighting their limitations in handling document diversity and ensuring fair evaluation. OmniDocBench establishes a robust, diverse, and fair evaluation standard for the document content extraction field, offering crucial insights for future advancements and fostering the development of document parsing technologies. The codes and dataset is available in https://github.com/opendatalab/OmniDocBench.
Problem

Research questions and friction points this paper is trying to address.

Evaluating document parsing methods lacks diversity and realism
Assessing performance across varied document types is limited
Current benchmarks lack comprehensive annotations and flexible evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Comprehensive annotations across diverse document types
Multi-level evaluations with 19 layout categories
Benchmarking pipeline-based and end-to-end models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Shanghai AI Laboratory | Abaka AI | 2077AI
L
Linke Ouyang
Shanghai AI Laboratory
Y
Yuan Qu
Shanghai AI Laboratory
Hongbin Zhou
Hongbin Zhou
Shanghai AI Laboratory
J
Jiawei Zhu
Shanghai AI Laboratory
R
Rui Zhang
Shanghai AI Laboratory
Qunshu Lin
Qunshu Lin
Co-Founder of Abaka.AI
Data-Centric AI
B
Bin Wang
Shanghai AI Laboratory
Z
Zhiyuan Zhao
Shanghai AI Laboratory
M
Man Jiang
Shanghai AI Laboratory
X
Xiaomeng Zhao
Shanghai AI Laboratory
J
Jin Shi
Shanghai AI Laboratory
F
Fan Wu
Shanghai AI Laboratory
P
Pei Chu
Shanghai AI Laboratory
M
Minghao Liu
2077AI
Z
Zhenxiang Li
Shanghai AI Laboratory
C
Chaoming Xu
Shanghai AI Laboratory
B
Bo Zhang
Shanghai AI Laboratory
Botian Shi
Botian Shi
Shanghai Artificial Intelligence Laboratory
VLMsDocument UnderstandingAutonomous Driving
Z
Zhongying Tu
Shanghai AI Laboratory
Conghui He
Conghui He
Shanghai AI Laboratory
Data-centric AILLMDocument Intelligence