WeVisDoc: From Coverage to Capability for Robust End-to-End Document Parsing

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出WeVisDoc框架,通过两阶段方法解决文档解析在不同布局和获取条件下的鲁棒性问题,提高了解析器的性能。
📝 Abstract
Document parsing converts document images into structured content and requires reliable performance across diverse layouts and acquisition conditions. Yet training corpora are biased toward common document types and clean digital pages, while expanding coverage alone does not specify how to address a parser's remaining weaknesses. We present WeVisDoc, a two-stage data-centric framework for robust end-to-end document parsing. Stage I broadens semantic, structural, and appearance coverage through heterogeneous data and structure-preserving degradation synthesis. Stage II uses a held-out probe to measure the Stage I parser's residual errors within fixed visual-structural clusters. These diagnostics guide targeted data construction and reallocation of the target-token budget. WeVisDoc-4B achieves an Overall score of 95.38 on OmniDocBench v1.6 and a mean Overall score of 75.54 across the three PureDocBench tracks, ranking first among the compared end-to-end parsers in all four settings. Compared with Stage I, Stage II improves Overall scores for the 2B and 4B models on both benchmarks, with larger gains on the degraded PureDocBench tracks, including a 4.03-point gain for the 4B model on the Real Degraded track.
Problem

Research questions and friction points this paper is trying to address.

document parsing
coverage
capability
robust
end-to-end
Innovation

Methods, ideas, or system contributions that make the work stand out.

two-stage data-centric framework
heterogeneous data
structure-preserving degradation synthesis
residual error measurement
targeted data construction
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Hao Yu
WeChat Vision, Tencent Inc.
K
Kang Liu
WeChat Vision, Tencent Inc.
L
Linnan Zhao
WeChat Vision, Tencent Inc.
J
Jiabo Zhan
WeChat Vision, Tencent Inc.
Chong Sun
Chong Sun
Tencent WeChat
Computer Vision
Chen Li
Chen Li
WeChat, Tencent
Computer Vision
Jing Lyu
Jing Lyu
Shanghai Jiao Tong University
Power electronicsstabilityrenewable energy grid integrationhigh-voltage dc transmission