TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing benchmarks in evaluating the recovery of heterogeneous evidence topologies within document layouts. To this end, it proposes the first layout-anchored multimodal GraphRAG benchmark. Methodologically, three controlled topologies are constructed in a bottom-up manner, and counterfactual verification is introduced to eliminate reasoning shortcuts, ensuring that models perform genuine reasoning by relying on cross-modal evidence. Additionally, a topology-aware evaluation metric system is designed. Empirical results demonstrate that while multimodal systems achieve optimal performance, notable deficiencies remain. The findings reveal the limitations of text-only and page-level visual retrieval in fine-grained topology recovery, establishing a new paradigm for topology-aware retrieval-augmented generation over complex documents.
📝 Abstract
Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at https://richardlrc.github.io/TopoGraphRAG-Bench/.
Problem

Research questions and friction points this paper is trying to address.

Multimodal GraphRAG
Evidence Topology
Layout-Grounded Reasoning
Document Understanding
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal GraphRAG
Layout-Grounded Benchmark
Evidence Topology
Counterfactual Validation
Topology-Aware Reasoning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Ruochi Li
Ruochi Li
North Carolina State University
Computer Science
J
Jianzhe Lin
Independent Researcher
Haoxuan Zhang
Haoxuan Zhang
UNIVERSITY OF NORTH TEXAS
Information scienceData scienceScientometrics
H
Haihua Chen
The Anuradha and Vikas Sinha Department of Data Science, University of North Texas, USA
J
Junhua Ding
School of Computing, University of Wyoming, USA
E
Edward Gehringer
Department of Computer Science, North Carolina State University, USA
Y
Yang Zhang
The Anuradha and Vikas Sinha Department of Data Science, University of North Texas, USA