DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning

📅 2026-08-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing approaches in long-document visual question answering, which often lack explicit modeling and traceability of reasoning evidence, thereby compromising answer accuracy and trustworthiness. The authors propose a hierarchical framework that formulates the task as explicit evidence graph reasoning, sequentially performing evidence localization, structured parsing, and graph-based multi-hop inference to enable traceable reasoning. A novel node-level evidence provenance mechanism is introduced, coupled with a two-stage training strategy combining supervised fine-tuning (SFT) and grouped relative policy optimization (GRPO) to jointly enhance localization and reasoning capabilities. The method outperforms the Qwen3-VL-8B-Instruct baseline by 14.4, 11.3, and 11.7 percentage points on MMLongBench-Doc, LongDocURL, and SlideVQA, respectively, while generating verifiable evidence graphs that significantly improve both performance and model transparency.
📝 Abstract
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded evidence is progressively composed during reasoning, limiting both answer accuracy and traceability. In this paper, we cast LongDocVQA as an explicit evidence graph reasoning problem rather than implicit answer prediction. To this end, we propose DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance. To effectively learn these capabilities, we develop a two-stage training framework: joint Supervised Fine-Tuning (SFT) first initializes evidence localization and graph reasoning abilities, followed by task-specific Group Relative Policy Optimization (GRPO) with dedicated rewards to further optimize these capabilities. Extensive experiments on MMLongBench-Doc, LongDocURL, and SlideVQA demonstrate that DocTrace consistently outperforms both existing open-source baselines and proprietary MLLMs. Compared with the Qwen3-VL-8B-Instruct backbone, DocTrace achieves absolute improvements of 14.4, 11.3, and 11.7 points on the three benchmarks, respectively. Beyond competitive performance, DocTrace constructs traceable evidence graphs with explicit node-level provenance, enabling transparent and verifiable reasoning for long document understanding.
Problem

Research questions and friction points this paper is trying to address.

Long Document Visual Question Answering
evidence traceability
multimodal reasoning
document understanding
explainable AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

evidence graph reasoning
hierarchical framework
traceable VQA
two-stage training
document parsing
L
Le Xiang
Baidu Basic Model Research and Development Department, Baidu Inc.
Z
Zhicheng Guan
Tsinghua Shenzhen International Graduate School, Tsinghua University
H
Hong Chen
Baidu Basic Model Research and Development Department, Baidu Inc.
X
Xiaocong Lin
Baidu Basic Model Research and Development Department, Baidu Inc.
Z
Zhenghua Lei
Baidu Basic Model Research and Development Department, Baidu Inc.
T
Teng Hu
Baidu Basic Model Research and Development Department, Baidu Inc.
B
Bolei He
Baidu Basic Model Research and Development Department, Baidu Inc.
Long Zeng
Long Zeng
Tsinghua University
Intelligent ManufacturingEmbodied AI RoboticsSketch Modeling