🤖 AI Summary
To address information fragmentation and cross-document reasoning challenges in complex retrieval tasks such as multi-hop question answering, this paper proposes HierRAG, a knowledge graph–driven hierarchical retrieval-augmented framework. HierRAG constructs a layered index graph integrating a knowledge graph layer and a collaborative document layer, leveraging graph neural networks to jointly model entity–document relationships—enabling coordinated coarse-grained semantic navigation and fine-grained knowledge localization. Unlike conventional flat RAG architectures, HierRAG introduces the first hierarchical indexing structure, significantly improving intra- and inter-document connectivity and multi-hop reasoning capability. Evaluated on five mainstream multi-hop QA benchmarks, HierRAG achieves substantial gains in both retrieval accuracy and response efficiency, demonstrating its effectiveness and generalizability in complex reasoning scenarios.
📝 Abstract
Large language models with retrieval-augmented generation encounter a pivotal challenge in intricate retrieval tasks, e.g., multi-hop question answering, which requires the model to navigate across multiple documents and generate comprehensive responses based on fragmented information. To tackle this challenge, we introduce a novel Knowledge Graph-based RAG framework with a hierarchical knowledge retriever, termed KG-Retriever. The retrieval indexing in KG-Retriever is constructed on a hierarchical index graph that consists of a knowledge graph layer and a collaborative document layer. The associative nature of graph structures is fully utilized to strengthen intra-document and inter-document connectivity, thereby fundamentally alleviating the information fragmentation problem and meanwhile improving the retrieval efficiency in cross-document retrieval of LLMs. With the coarse-grained collaborative information from neighboring documents and concise information from the knowledge graph, KG-Retriever achieves marked improvements on five public QA datasets, showing the effectiveness and efficiency of our proposed RAG framework.