PILAR: A Page-Grounded Unified Evidence Representation via an Entity-Linked Assertion Graph for Open-Domain QA Agents over Multimodal Document Corpora

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of cross-modal linking caused by dispersed evidence in multimodal documents. To this end, it proposes a page-anchored unified assertion space mapping mechanism that integrates textual, tabular, and visual facts into an entity-linking assertion graph. This approach enables effective association of cross-modal evidence, substantially enhancing retrieval-augmented question answering for compositional and multi-hop queries. Experimental results demonstrate that the proposed framework achieves state-of-the-art end-to-end Exact Match (EM) and Average Normalized Levenshtein Similarity (ANLS) scores across multiple benchmarks. Notably, it yields a 5.9-point EM improvement on three-hop questions, validating the effectiveness of the unified representation for complex multimodal reasoning.
📝 Abstract
Open-domain question answering (ODQA) over multimodal document corpora requires linking evidence scattered across text, tables, and figures. Existing systems often store these sources separately or retrieve only coarse pages, which weakens global evidence linking. We present PILAR, a page-grounded unified evidence representation instantiated as an entity-linked assertion graph. PILAR maps sentence-, table-, and figure-derived facts into a common assertion space and uses the graph as a controlled linking layer over robust page retrieval. In a shared-reader evaluation with four agent frameworks, fourteen retrieval backends, and two benchmarks, PILAR achieves the best end-to-end EM/ANLS. Gains are largest on compositional, cross-document, and multimodal questions, with a single-shot improvement of +1.6 EM over flat retrieval, rising to +2.9 on compositional and +5.9 on 3-hop questions. Ablations show that current gains are driven mainly by the text-instantiated slice of the framework, while visual assertions help only after locality-aware filtering. We therefore position PILAR as a unified evidence representation for multimodal ODQA rather than a standalone visual-reasoning module.
Problem

Research questions and friction points this paper is trying to address.

Open-domain question answering
Multimodal document corpora
Evidence linking
Cross-document reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Open-Domain Question Answering
Assertion Graph
Multimodal Document Corpora
Entity Linking
Unified Evidence Representation
J
Joongmin Shin
Korea University
G
Gyuho Shim
Korea University
J
Jung-hun Lee
Korea Maritime and Ocean University
Jaehyung Seo
Jaehyung Seo
Korea University
Natural Language GenerationCommonsense ReasoningHallucinationKnowledge Editing