PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing document parsing approaches: end-to-end methods suffer from excessively long decoding paths due to serialized layout and content generation, while two-stage methods enable parallelism at the cost of redundant feature extraction and disrupted page-level context. To overcome these issues, the authors propose PaDoc, the first framework to achieve layout-guided parallel decoding within a single large multimodal language model. PaDoc constructs a branching structure based on predicted layouts and generates region-specific content in parallel over a shared page representation, preserving full contextual information. Leveraging techniques such as prefix-conditioned factorization, packed variable-length ancestor attention, and masked parallel decoding, PaDoc enables efficient inference on the vLLM backend. It achieves state-of-the-art performance on OmniDocBench Full with a layout F1 of 91.1, an overall score of 94.24, and a text edit error rate of only 0.038, while improving throughput by 67.4–118% and reducing P95 latency by 39.2–54.9% over baselines.
📝 Abstract
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total content, whereas crop-based two-stage parsers expose region-level parallelism at the cost of repeated visual prefills and fragmented page context. To retain full-page context while removing dependencies, we propose PaDoc, a layout-grounded parser that treats the predicted layout as a branching structure over a shared page representation. Under a region-sufficiency assumption, we derive a prefix-conditioned factorization in which the layout stream and regional content branches advance concurrently, reducing the decoding depth to the longest layout-content path. We realize this factorization within a single MLLM: packed variable-length ancestor attention preserves the visibility under standard next-token training, while masked parallel decoding creates branches that the evaluated vLLM backend serves as concurrent requests with cache-resident shared-prefix reuse. On OmniDocBench Full, PaDoc attains an Overall layout F1 of 91.1 and, among end-to-end parsers, a top-tier Overall score of 94.24 together with the best Text Edit (0.038) and Formula CDM (95.59). On a 384-page subset and one A800 GPU, it is the fastest end-to-end parser at five concurrency levels, improving valid-page throughput by 67.4-118% and reducing P95 latency by 39.2-54.9% relative to a same-backbone Sequential SFT baseline. Code is available at https://github.com/Longin-Yu/Padoc
Problem

Research questions and friction points this paper is trying to address.

document parsing
layout grounding
parallel decoding
end-to-end parsing
page context
Innovation

Methods, ideas, or system contributions that make the work stand out.

parallel decoding
layout-grounded parsing
shared-prefix reuse
end-to-end document parsing
region-sufficiency assumption
🔎 Similar Papers
No similar papers found.