Score
Design and implement retrieval-augmented generation systems that represent knowledge as multimodal graphs whose nodes and edges encode text, images, and other modalities, and that construct, index, and retrieve relevant subgraphs for a given query. Build methods to fuse retrieved multimodal evidence into prompts or structured contexts for large models and to perform graph-based multi‑hop reasoning across pages or documents to answer document‑level visual and cross‑modal queries.
Large language models (LLMs) suffer from factual hallucinations due to outdated knowledge and training data limitations. While retrieval-augmented generation (RAG) mitigates this issue, the multifaceted roles of graph technologies in RAG remain unstructured and lack a unified framework. This paper proposes the first taxonomy of RAG grounded in “graph functionality,” systematically characterizing graphs’ distinct roles in knowledge organization, semantic retrieval, relational reasoning, and dynamic knowledge updating. We integrate techniques—including knowledge graph embedding, subgraph retrieval, graph neural networks, and graph database optimization—to enable efficient ingestion of structured and semi-structured knowledge. Through a comprehensive analysis of 120+ works, we identify critical bottlenecks such as graph sparsity modeling and real-time update latency, and propose six future directions—including scalable graph indexing and causality-aware retrieval—to bridge interdisciplinary gaps across graph learning, databases, and NLP.
Large language models (LLMs) face critical bottlenecks in specialized domains, including weak comprehension of complex problems, difficulty integrating cross-source knowledge, and low information processing efficiency. To address these challenges, this paper presents a systematic review of the Graph-Augmented Retrieval-Augmented Generation (GraphRAG) paradigm and introduces three core innovations: (1) domain-knowledge-guided graph-structured representation, explicitly modeling entity relationships and hierarchical semantics; (2) graph neural network–based retrieval supporting multi-hop reasoning; and (3) structure-aware knowledge fusion and logically consistent generation. The approach integrates knowledge graph construction, structured prompt engineering, and controllable generation techniques. We open-source the first comprehensive GraphRAG repository on GitHub, featuring multi-domain implementation cases, a clear taxonomy of key technical challenges, and a roadmap for methodological evolution. This work provides a principled, end-to-end framework and practical benchmark for deploying domain-specialized LLMs.
Existing RAG frameworks are largely confined to text modality, struggling to effectively retrieve and reason over real-world documents containing multimodal elements such as images, tables, and mathematical formulas. This work introduces MultiRAG—the first RAG framework supporting unified retrieval and generation across full modalities (text, image, table, formula). Its core innovation lies in constructing a dual-graph architecture: a structural graph modeling document layout and entity relations, and a semantic graph encoding cross-modal semantics; together with a hybrid retrieval mechanism integrating structural navigation and semantic matching. By representing multimodal entities explicitly and aligning them into a unified embedding space, MultiRAG enables fine-grained knowledge navigation and cross-modal semantic matching. It achieves significant improvements over state-of-the-art methods on multimodal benchmarks, especially for long-document scenarios. The code and models are publicly released, advancing RAG toward a truly multimodal paradigm.
To address three key challenges in audio-visual multimodal RAG—narrow modality coverage of knowledge graphs, weak multi-hop connectivity, and imprecise retrieval—this paper proposes a query-aligned Multi-hop Multimodal Knowledge Graph (M³KG) construction and retrieval framework. We introduce a novel lightweight multi-agent construction method that significantly expands the modality granularity and cross-modal path depth of multimodal knowledge graphs (MMKGs). Furthermore, we design the GRASP mechanism—comprising query-driven entity anchoring, supportiveness assessment, and redundant context pruning—to enhance retrieval precision. By integrating modality-aware retrieval, query grounding, relevance scoring, and embedding alignment, our approach improves fact consistency and cross-modal localization accuracy for multimodal large language models (MLLMs) in multi-hop reasoning. Extensive evaluation across multiple multimodal benchmarks demonstrates substantial gains in answer faithfulness and reasoning depth.
Existing RAG systems are limited to unimodal text retrieval and struggle to effectively process unstructured multimodal documents containing text, images, tables, mathematical formulas, and charts. To address this, we propose MAHA—a modality-aware hybrid retrieval architecture that innovatively integrates dense vector retrieval with knowledge graph–driven structured traversal. First, MAHA constructs a modality-aware knowledge graph that explicitly encodes cross-modal semantic relationships. Then, it jointly performs question answering and interpretable reasoning via synergistic multimodal embedding and graph traversal. Evaluated on multiple benchmark datasets, MAHA achieves state-of-the-art performance (e.g., ROUGE-L = 0.486), significantly outperforming baseline methods. It is the first approach to achieve full-modality coverage—supporting text, images, tables, formulas, and diagrams—while delivering both high retrieval coverage and strong interpretability. Comprehensive experiments validate MAHA’s effectiveness, robustness, and scalability across diverse multimodal document understanding tasks.
To address the limitation of conventional RAG methods—which retrieve isolated text snippets while ignoring the inherent topological structure of networked documents (e.g., citation graphs, knowledge graphs)—this paper proposes a graph-aware RAG framework. Methodologically: (1) we design a divide-and-conquer linear-time text subgraph retrieval algorithm enabling efficient subgraph-level retrieval; (2) we introduce a dual-path encoder jointly processing text and graph views to explicitly model structural relationships; and (3) we incorporate topology-aware prompt injection and multi-hop graph reasoning fine-tuning. Evaluated on multiple graph reasoning benchmarks, our approach significantly outperforms existing RAG methods, achieving a 21.4% absolute accuracy gain on complex 3+-hop reasoning tasks. To our knowledge, this is the first work to synergistically enhance generative outputs through joint optimization of textual semantics and graph topology.
This work addresses the limitations of existing graph-augmented retrieval methods, which treat knowledge graphs as static structures and thus struggle to support cross-modal, iterative, and revisable reasoning. The paper proposes a self-evolving graph retrieval framework that unifies knowledge graph construction and retrieval within an agent-driven, closed-loop evolutionary mechanism. By modeling multimodal knowledge as a dynamic hypergraph environment, the framework enables an intelligent agent to perform adaptive multi-hop reasoning through actions such as querying, expanding, editing, and answering. The agent’s behavior is governed by a Markov decision process, allowing it to learn optimal reasoning strategies. Evaluated on multimodal visual and textual question-answering benchmarks, the approach significantly outperforms current retrieval-augmented generation (RAG) methods, achieving notable advances in accuracy, knowledge coverage, and traceable reasoning.
This work addresses the limitations of existing retrieval-augmented generation (RAG) systems in supporting complex cross-modal reasoning within multimodal large language models, where conventional graph construction relies on costly text translation and often discards fine-grained visual information. To overcome these challenges, the authors propose a lightweight, multi-granularity graph RAG framework that unifies textual entities and visual regions into cohesive multimodal nodes. The approach leverages lightweight text parsing and entity-driven visual grounding to construct a hierarchical multimodal knowledge graph, complemented by a multi-granularity graph retrieval mechanism enabling structured multi-hop reasoning. Evaluated across four multimodal benchmarks, the method achieves state-of-the-art performance while accelerating graph construction by 43.3× and reducing associated costs by 23.9×.
Document understanding faces three key challenges: (1) OCR systems discard structural information; (2) multimodal large language models (MLLMs) exhibit limited contextual modeling capacity; and (3) conventional RAG frameworks struggle with heterogeneous modalities—text, tables, charts, and layout. To address these, this paper proposes a comprehensive multimodal RAG framework for document understanding. Methodologically, it introduces a novel “domain–modality–granularity” three-dimensional taxonomy and integrates OCR, MLLMs, vector retrieval, graph neural networks, and agent-based coordination to enable cross-modal joint modeling and reasoning. It further incorporates graph-structured document representation and an agent-driven retrieval–reasoning closed loop, substantially improving fine-grained comprehension, retrieval efficiency, and system robustness. The work also surveys mainstream datasets, evaluation benchmarks, and real-world applications, identifies open challenges, and provides a holistic technical roadmap for future research.
Effectively retrieving, reasoning, and understanding multimodal information remains a critical challenge for agentic systems. Traditional Retrieval-augmented Generation (RAG) methods rely on linear interaction histories, which struggle to handle long-context tasks, especially those involving information-sparse yet token-heavy visual data in iterative reasoning scenarios. To bridge this gap, we introduce VimRAG, a framework tailored for multimodal Retrieval-augmented Reasoning across text, images, and videos. Inspired by our systematic study, we model the reasoning process as a dynamic directed acyclic graph that structures the agent states and retrieved multimodal evidence. Building upon this structured memory, we introduce a Graph-Modulated Visual Memory Encoding mechanism, with which the significance of memory nodes is evaluated via their topological position, allowing the model to dynamically allocate high-resolution tokens to pivotal evidence while compressing or discarding trivial clues. To implement this paradigm, we propose a Graph-Guided Policy Optimization strategy. This strategy disentangles step-wise validity from trajectory-level rewards by pruning memory nodes associated with redundant actions, thereby facilitating fine-grained credit assignment. Extensive experiments demonstrate that VimRAG consistently achieves state-of-the-art performance on diverse multimodal RAG benchmarks. The code is available at https://github.com/Alibaba-NLP/VRAG.