Signal or Noise? Modality Contribution and Cooperation in Multimodal GraphRAG

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of redundant evidence interference and unclear modality contribution mechanisms in multimodal GraphRAG. Focusing on the DocVQA scenario, it proposes an edge-level modality filtering mechanism grounded in multimodal knowledge graphs, enabling modality-aware retrieval through graph-structural pruning. The research reveals that inter-modal interactions are predominantly characterized by redundancy rather than synergy, with text and tables providing the strongest contributions, while non-textual modalities exhibit positive collaborative effects. Experimental results demonstrate that selective retrieval significantly outperforms unified retrieval approaches. By elucidating these modality dynamics and validating the efficacy of structure-guided filtering, this work establishes a novel paradigm for optimizing reasoning processes within multimodal GraphRAG systems.
📝 Abstract
Multimodal knowledge graphs (KGs) integrate information from text, figures, tables, and other modalities into a unified structured representation, with the promise that richer evidence enables better inference. In GraphRAG systems built over such graphs, it is commonly assumed that retrieving evidence from more modalities at inference time improves downstream performance. Yet, redundant or overlapping multimodal evidence may distract language models in question answering (QA), and whether each modality contributes equally across questions, models, and tasks remains poorly understood. In this work, we study how modality-aware retrieval affects downstream inference in a multimodal GraphRAG pipeline, using document visual question answering (DocVQA) as a testbed. We extend an existing KG-based QA framework to be modality-aware, leveraging the graph structure to track which modality supports which facts and to selectively filter evidence at the edge level. This enables us to investigate whether providing all available multimodal evidence at inference time benefits QA, and to evaluate the contribution and cooperation of modalities across question, task, and model characteristics. Through a controlled analysis within a state-of-the-art multimodal GraphRAG pipeline, five multimodal LLMs and two DocVQA benchmarks, we find that tables and text provide the strongest contributions, and that combining modalities frequently produces redundancy rather than synergy, particularly for pairs involving textual information. Positive cooperation appears mainly between non-text modalities and depends on question intent and task type. Our findings argue for selective, modality-aware retrieval in the design of more effective GraphRAG systems, where modalities are filtered according to the downstream task rather than retrieved uniformly.
Problem

Research questions and friction points this paper is trying to address.

Multimodal GraphRAG
Modality Contribution
Modality Cooperation
Document Visual Question Answering
Evidence Redundancy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal GraphRAG
Modality-aware Retrieval
Document Visual Question Answering
Edge-level Filtering
Modality Cooperation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.