Score
Designs and implements parsers and processors for tree‑structured markup (XML/DOM) and for URLs, including tokenization, syntax validation, and construction of document object trees. Builds extractors, normalizers, link resolvers, and transformers that fetch and parse document sources, traverse and identify regions of interest, canonicalize markup and links, and produce structured content and metadata for downstream use.
This work addresses high-accuracy conversion of unstructured/semi-structured documents (e.g., contracts, academic papers, invoices) into structured, machine-readable data. Method: We systematically survey and empirically compare modular pipeline approaches against end-to-end multimodal large models, proposing a unified framework integrating OCR, layout analysis (LayoutParser), graph neural networks, vision-language models (VLMs), and specialized formula/table recognition. We identify and characterize core bottlenecks—layout understanding, dense text recognition, and cross-modal alignment—for the first time. Contribution/Results: We establish a comprehensive analytical framework covering methodology, challenges, and benchmarks, revealing >32% performance gaps of current SOTA on complex layouts (e.g., multi-column, nested tables). We propose a “dual-driven” evolution path emphasizing both data diversity and scale, and open-source a larger annotated dataset to significantly advance knowledge base construction and training-data generation for large models.
This work addresses the challenges of structured information extraction from web pages, where handcrafted rules are brittle and large language models incur prohibitive computational costs. The authors propose a novel paradigm that treats the HTML DOM as a prunable tree, employing a dedicated pruning mechanism to retain high-information-density subtrees. A lightweight 0.6B-parameter language model is then applied for zero-shot cross-domain extraction. To ensure result reliability, the approach incorporates a traceable Grounded XPath Resolution (GXR) mechanism. Evaluated on the SWDE dataset, the method achieves an F1 score of 88.1%, outperforming several larger, fully trained baseline models while substantially reducing computational and deployment overhead.
This work addresses severe parsing errors in dense document pages caused by unstable layout assumptions, which lead to mismatches between detector outputs and the input sequence expected by the parser. To resolve this, the authors introduce a lightweight structural refinement module between a DETR-style detector and the parser, performing set-level reasoning using query features, semantic cues, bounding box geometry, and visual evidence. This module jointly decides which instances to retain, refines bounding boxes, and predicts the correct parsing order. By integrating retention-oriented supervision with a difficulty-aware ordering objective, the method significantly enhances layout–parsing interface consistency under complex layouts, reducing the reading order edit distance to 0.024 on OmniDocBench while consistently improving page-level layout quality.
This work addresses the challenges of fine-grained product information extraction from shopping review webpages and the difficulty of dynamically updating product databases. To this end, we propose MarkupLM++, the first model to extend sequence labeling to internal nodes—not only leaf nodes—of the DOM tree, thereby explicitly modeling hierarchical webpage structure and semantic associations. Built upon the MarkupLM architecture, MarkupLM++ incorporates large-scale, diverse web annotation data and explicit DOM tree structure for end-to-end fine-tuning. Experiments on a real-world e-commerce review dataset yield 90.6% precision, 72.4% recall, and 80.5% F1-score, significantly outperforming baseline methods. This work advances web document understanding from “textual content extraction” toward “structured semantic parsing,” establishing a new paradigm for constructing real-time-updating product knowledge bases and enabling downstream applications such as customer analytics and recommendation.
This work addresses the limitation of existing document analysis systems that flatten complex documents into plain text, thereby discarding critical hierarchical structures such as sections, tables, and figures, which hinders effective filtering and in-depth analysis. To overcome this, the authors propose a structure-aware document understanding framework that integrates full document hierarchy into semantic indexing and analysis. The approach constructs a hierarchical document tree by parsing the original layout, leverages large language models to generate structure-aware semantic representations, and introduces a multi-view interactive web interface enabling precise natural language–driven retrieval and question answering. Experiments demonstrate significant improvements in retrieval accuracy and question-answering performance on diverse complex documents, including academic papers, technical manuals, and financial reports. The code and a live demo system are publicly released.
This work addresses the limitations of existing methods in question answering over semi-structured documents containing tables, figures, and hierarchical paragraphs, which often suffer from fragmented OCR outputs, inadequate hierarchical modeling, and insufficient cross-region alignment. To overcome these challenges, we propose MoDora, a novel system that reconstructs OCR results through local alignment aggregation, explicitly models inter-component hierarchical and spatial relationships via a Component Correlation Tree (CCTree), and introduces a question-type-aware hybrid retrieval mechanism that fuses semantic and layout information. Evaluated on multiple benchmarks, MoDora significantly outperforms current state-of-the-art approaches, achieving absolute accuracy gains of 5.97% to 61.07% and enabling precise, layout- and semantics-driven document understanding.
This study addresses the lack of systematic evaluation on how PDF preprocessing frameworks impact downstream domain-specific question answering. For the first time, it directly links PDF conversion quality to RAG-based QA performance by constructing a benchmark using Portuguese administrative documents. The work systematically compares four open-source tools—Docling, MinerU, Marker, and DeepSeek OCR—across various text extraction, cleaning, chunking, and metadata strategies, and includes GraphRAG for comparative analysis. Results demonstrate that hierarchical chunking and metadata enrichment significantly outperform the choice of conversion tool itself in boosting QA accuracy. The optimal configuration (Docling with hierarchical chunking and image descriptions) achieves 94.1% accuracy—approaching human-annotated performance (97.1%)—whereas GraphRAG attains only 82%, underscoring the critical role of structured preprocessing.
Traditional approaches rely on manually designed annotation schemas and exhaustive document labeling, which are costly and difficult to scale. This work proposes an end-to-end framework leveraging large language models to automatically transform natural language research questions and raw text into structured databases, supported by an interactive interface that enables user-guided refinement. The method establishes, for the first time, a closed-loop pipeline from research questions to structured evidence, integrating expert feedback and domain-adaptation mechanisms. Evaluated in legal and computational biology domains, it significantly enhances the efficiency and accuracy of cross-domain information extraction. The system, along with its public web interface, has been open-sourced.
This work addresses the limitations of existing unified document parsing models, which rely on full-page autoregressive generation and thus struggle to scale to long documents while failing to exploit global layout structure and parallelism across content blocks. To overcome these challenges, the authors propose a hierarchical parallel decoding paradigm that employs a master layout branch to orchestrate document structure and dynamically allocates content blocks to concurrent decoding branches. This approach is further enhanced by progressive multi-token prediction (P-MTP) to reduce decoding steps. The method represents the first integration of hierarchical parallelism into unified document parsing, achieving a throughput of 4,752 tokens/s on public benchmarks—2.62× faster than the current fastest model and 3.06× faster than a standard autoregressive baseline—while maintaining competitive accuracy.