Score
Designs and implements retrieval systems that index and search nodes, paths, and subgraphs in knowledge graphs (including temporal KGs), using techniques such as graph-based indexing, random-walk or walk-based retrieval, path-finding, and temporal filtering to fetch relevant event chains or multi-hop evidence. Builds ranking and selection components that order nodes/paths/subgraphs by structural relevance or centrality and produce evidence subgraphs suitable for downstream consumers (e.g., LLM augmentation).
Large language models (LLMs) face critical bottlenecks in specialized domains, including weak comprehension of complex problems, difficulty integrating cross-source knowledge, and low information processing efficiency. To address these challenges, this paper presents a systematic review of the Graph-Augmented Retrieval-Augmented Generation (GraphRAG) paradigm and introduces three core innovations: (1) domain-knowledge-guided graph-structured representation, explicitly modeling entity relationships and hierarchical semantics; (2) graph neural network–based retrieval supporting multi-hop reasoning; and (3) structure-aware knowledge fusion and logically consistent generation. The approach integrates knowledge graph construction, structured prompt engineering, and controllable generation techniques. We open-source the first comprehensive GraphRAG repository on GitHub, featuring multi-domain implementation cases, a clear taxonomy of key technical challenges, and a roadmap for methodological evolution. This work provides a principled, end-to-end framework and practical benchmark for deploying domain-specialized LLMs.
Existing retrieval methods for text-rich graph knowledge bases (TG-KBs) suffer from a disconnect between structural and textual knowledge access; hybrid approaches either neglect structural retrieval or underutilize neighborhood interactions. Method: We propose the Planning–Reasoning–Organizing (PRO) framework—the first to integrate *textualized planning graphs* into retrieval. PRO employs query-driven structural traversal planning, jointly performs graph path reasoning and fine-grained text matching, and leverages structural trajectories to enhance candidate re-ranking—enabling dynamic cross-modal alignment and complementary knowledge enhancement. The method unifies interpretable planning graph generation, multi-stage neural ranking, and semantic-structural joint modeling. Contribution/Results: PRO achieves significant improvements over state-of-the-art methods across multiple TG-KB benchmarks, empirically validating both the effectiveness of structural trajectory modeling for re-ranking and the robustness and generalizability of the hybrid retrieval paradigm.
To address information fragmentation and cross-document reasoning challenges in complex retrieval tasks such as multi-hop question answering, this paper proposes HierRAG, a knowledge graph–driven hierarchical retrieval-augmented framework. HierRAG constructs a layered index graph integrating a knowledge graph layer and a collaborative document layer, leveraging graph neural networks to jointly model entity–document relationships—enabling coordinated coarse-grained semantic navigation and fine-grained knowledge localization. Unlike conventional flat RAG architectures, HierRAG introduces the first hierarchical indexing structure, significantly improving intra- and inter-document connectivity and multi-hop reasoning capability. Evaluated on five mainstream multi-hop QA benchmarks, HierRAG achieves substantial gains in both retrieval accuracy and response efficiency, demonstrating its effectiveness and generalizability in complex reasoning scenarios.
This work addresses the semantic mismatch caused by structural heterogeneity in knowledge graph question answering and the lack of global structural awareness in existing methods. It reframes multi-hop reasoning as a schema-guided graph search task, constructing an adaptive query schema graph via a semantic-structure projection mechanism. By integrating a Triple-Dependent Graph Neural Network, the approach enables globally guided node anchoring and subgraph retrieval, thereby incorporating global structural information into the retrieval phase for the first time to generate high-quality evidence reasoning graphs. This strategy substantially improves both retrieval accuracy and evidence completeness of multi-hop reasoning paths, achieving state-of-the-art performance across multiple benchmark datasets.
This work addresses the limitations of existing retrieval-augmented generation methods, which struggle to simultaneously support structured constraints and multi-hop reasoning, as well as knowledge graph–based approaches that suffer from semantic fragmentation, high maintenance costs, and difficulty in updating. The authors propose the SAG architecture, which organizes documents via an event–entity index, preserving n-ary relations within original text blocks without constructing a global knowledge graph. At query time, SAG dynamically connects relevant event blocks through shared entities to form an evidence neighborhood. A novel dynamic hyperedge mechanism is introduced to avoid decomposing relations into triples, thereby maintaining semantic integrity while enabling efficient incremental updates and complex reasoning. By integrating SQL-style structured retrieval with dynamic connection, SAG achieves state-of-the-art performance on HotpotQA, 2WikiMultiHopQA, and MuSiQue, attaining a Recall@5 of 80.36% on MuSiQue—11.52 percentage points above the strongest baseline.
To address the challenge of simultaneously achieving high retrieval efficiency, accurate reasoning, and effective hallucination suppression when augmenting large language models (LLMs) with knowledge graphs (KGs), this paper proposes SubgraphRAG—a fine-tuning-free, dynamic subgraph retrieval framework. Its core innovation lies in a lightweight MLP coupled with a parallel triple-scoring mechanism that explicitly encodes directed structural distances within the KG. Crucially, SubgraphRAG adaptively adjusts subgraph size based on query difficulty and LLM capability, enabling joint optimization of retrieval granularity and reasoning efficacy. The method integrates KG subgraph retrieval, structure-aware scoring, and LLM-coordinated generation. Evaluated on WebQSP and ComplexWebQuestions (CWQ), SubgraphRAG significantly reduces hallucination rates while improving answer accuracy and interpretability. It achieves state-of-the-art performance with Llama3.1-8B-Instruct and sets new benchmark records with GPT-4o on both datasets.
This work addresses the efficiency and scalability challenges posed by the irregular structure of knowledge graphs in multi-hop compositional query answering. The authors propose SeedER, a novel framework that uniquely integrates a seed-expansion mechanism with reinforcement learning. It first generates compact sets of core seed entities through lightweight dense and sparse retrieval, then iteratively expands these seeds via a graph-aware policy, decomposing global reasoning into reusable local decisions at low computational cost. By maintaining concise candidate sets while substantially improving recall, SeedER outperforms strong baselines and functions as an efficient single-stage retriever. The approach also offers theoretical advantages in compositional generalization and submodular optimization under graph constraints.
Traditional vector retrieval struggles to handle complex queries requiring structured reasoning in industrial knowledge graphs. This work constructs an aerospace supply chain knowledge graph comprising 46 node types and 64 relation types, and introduces the “operator vocabulary” hypothesis, positing that the bottleneck in graph reasoning lies not in model intelligence but in the availability of computational primitives. Guided by this insight, the authors design an LLM-driven query planner that integrates nine graph traversal primitives and six graph computation tools, forming a structure-aware retrieval-augmented generation framework. Evaluated on 23 queries spanning ten intent categories, the approach achieves an F1 score of 0.632, significantly outperforming a customized processor (0.472), and further exposes a systematic bias in existing entity-level F1 metrics when assessing structured queries.
Existing knowledge graph multi-hop retrieval systems struggle to simultaneously achieve efficiency, scalability, and interpretability. This work proposes a hardware-aligned symbolic retrieval framework that reformulates multi-hop reasoning as efficient hardware operations over decomposed representations of subjects, predicates, and objects. By integrating degree-aware graph partitioning, cross-partition routing, and on-demand caching, the approach enables highly efficient retrieval at billion-edge scale. It is the first method to deliver interpretable, scalable, and hardware-efficient multi-hop retrieval on both CPUs and GPUs, significantly accelerating inference while preserving high fidelity. The framework has been successfully deployed in biomedical domains for collaborative reasoning between knowledge graphs and large language models.
Traditional semantic search struggles to model the hierarchical structures and multi-hop cross-references prevalent in enterprise documents, limiting retrieval accuracy. This work proposes an agent-driven, recursive knowledge graph construction approach that automatically parses substitutional logic and cross-level references among documents to generate a structured graph representation. The resulting knowledge graph is integrated into a retrieval-augmented generation (RAG) framework to enable precise querying of complex regulatory logic. Evaluated on the Code of Federal Regulations benchmark, the proposed method achieves a 70% improvement in question-answering accuracy over standard vector-based RAG systems, substantially overcoming the limitations of conventional semantic retrieval.