knowledge-graph augmented retrieval

Designs and implements retrieval and search systems that use structured knowledge graphs to index, link, expand, and rank candidate items by mapping queries and documents to entities and traversing graph relations. Builds components for entity-linked and relation-aware nearest-neighbor retrieval, multi-hop evidence retrieval, KG-augmented prediction and generation (graph-RAG), ontology-backed indexing and disambiguation, and provenance-aware or cross-lingual source linking.

knowledge-graphaugmentedretrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Graph-based Approaches and Functionalities in Retrieval-Augmented Generation: A Comprehensive Survey

Apr 08, 2025
ZZ
Zulun Zhu
🏛️ Nanyang Technological University | Beijing University of Posts and Telecommunications

Large language models (LLMs) suffer from factual hallucinations due to outdated knowledge and training data limitations. While retrieval-augmented generation (RAG) mitigates this issue, the multifaceted roles of graph technologies in RAG remain unstructured and lack a unified framework. This paper proposes the first taxonomy of RAG grounded in “graph functionality,” systematically characterizing graphs’ distinct roles in knowledge organization, semantic retrieval, relational reasoning, and dynamic knowledge updating. We integrate techniques—including knowledge graph embedding, subgraph retrieval, graph neural networks, and graph database optimization—to enable efficient ingestion of structured and semi-structured knowledge. Through a comprehensive analysis of 120+ works, we identify critical bottlenecks such as graph sparsity modeling and real-time update latency, and propose six future directions—including scalable graph indexing and causality-aware retrieval—to bridge interdisciplinary gaps across graph learning, databases, and NLP.

Addressing factual errors in LLMs via graph-based RAGIdentifying challenges and future directions in graph-enhanced RAGSurveying graph roles in RAG for structured knowledge

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge that existing retrieval-augmented generation systems struggle to effectively aggregate dispersed evidence from multiple sources when handling complex queries, while approaches relying on explicit knowledge graphs suffer from high construction costs and poor compatibility. To overcome these limitations, the authors propose a graph-structured augmentation and reranking method that avoids building a full knowledge graph. During offline preprocessing, data objects are enriched with graph-based contextual information; at inference time, candidate results are reranked using graph-aware proximity measures. The approach is retrieval-agnostic, seamlessly integrates with mainstream vector databases, and significantly improves retrieval performance across multiple benchmarks—while maintaining low inference latency and strong system compatibility.

complex information needsinformation retrievalknowledge graph

This work proposes an end-to-end graph-based retrieval-augmented generation (RAG) framework that addresses the limitations of traditional RAG methods in efficiently retrieving relevant information within unknown search spaces or when handling semi-structured and structured documents. By integrating labeled property graphs (LPGs) with the Resource Description Framework (RDF), the approach automatically converts JSON key-value pairs into RDF triples to incorporate semi-structured data. It further introduces a text-to-Cypher query generation mechanism, enabling real-time, high-precision graph retrieval without requiring a predefined number of source documents. Eliminating inefficient re-ranking steps, the method significantly enhances answer accuracy, reasoning capability, and overall response quality, demonstrating particularly strong performance in complex semi-structured tasks and online scenarios.

knowledge-intensive tasksRetrieval-Augmented Generationsemi-structured data

To address the high computational cost, graph retrieval latency, and scalability bottlenecks of GraphRAG systems in enterprise-scale unstructured text, this paper proposes a lightweight, LLM-free graph-augmented generation framework. Methodologically, it introduces a dependency-parsing–based knowledge graph construction pipeline leveraging industrial-grade NLP tools for efficient entity and relation extraction; further, it employs hybrid query node identification and single-hop traversal to enable low-latency subgraph retrieval. The key innovation lies in decoupling both graph construction and retrieval from LLM dependence while preserving multi-hop reasoning capability and achieving high recall. Evaluated on the SAP dataset, the framework improves LLM-as-Judge evaluation scores by 15% over conventional RAG, reduces graph construction cost significantly, and attains 94% of the retrieval performance of LLM-based baselines.

High computational cost of LLM-based knowledge graph constructionLatency issues in graph-based retrieval systemsScalability challenges in enterprise GraphRAG deployment

Existing GraphRAG approaches for knowledge graph (KG) question answering over graph databases commonly neglect or underutilize the retrieval step, leading to inaccurate Cypher query generation and frequent hallucinations. To address this, we propose the first plug-and-play end-to-end framework that deeply integrates retrieval-augmented generation (RAG) with large language model (LLM) fine-tuning—enabling precise multi-hop Cypher query generation and verifiable reasoning. Our method unifies subgraph context construction, native graph database interfacing, LLM fine-tuning, and retrieval-augmented inference. Evaluated on two major text-attribute KG QA benchmarks, our approach consistently outperforms state-of-the-art methods across all four metrics. It further achieves high sample efficiency during training and strong system scalability, making it both practically deployable and theoretically grounded.

Enhances accuracy in graph database QAGenerates correct Cypher queries for subgraphsImproves retrieval for multi-hop KG queries

KG-Retriever: Efficient Knowledge Indexing for Retrieval-Augmented Large Language Models

Dec 07, 2024
WC
Weijie Chen
🏛️ Beijing University of Posts and Telecommunications | University of Science and Technology Beijing | Xiaomi Corporation

To address information fragmentation and cross-document reasoning challenges in complex retrieval tasks such as multi-hop question answering, this paper proposes HierRAG, a knowledge graph–driven hierarchical retrieval-augmented framework. HierRAG constructs a layered index graph integrating a knowledge graph layer and a collaborative document layer, leveraging graph neural networks to jointly model entity–document relationships—enabling coordinated coarse-grained semantic navigation and fine-grained knowledge localization. Unlike conventional flat RAG architectures, HierRAG introduces the first hierarchical indexing structure, significantly improving intra- and inter-document connectivity and multi-hop reasoning capability. Evaluated on five mainstream multi-hop QA benchmarks, HierRAG achieves substantial gains in both retrieval accuracy and response efficiency, demonstrating its effectiveness and generalizability in complex reasoning scenarios.

Addresses information fragmentation via hierarchical knowledge indexingEnhances multi-hop question answering in retrieval-augmented LLMsImproves cross-document retrieval efficiency using graph structures

Latest Papers

What's happening recently
View more

Traditional vector retrieval struggles to handle complex queries requiring structured reasoning in industrial knowledge graphs. This work constructs an aerospace supply chain knowledge graph comprising 46 node types and 64 relation types, and introduces the “operator vocabulary” hypothesis, positing that the bottleneck in graph reasoning lies not in model intelligence but in the availability of computational primitives. Guided by this insight, the authors design an LLM-driven query planner that integrates nine graph traversal primitives and six graph computation tools, forming a structure-aware retrieval-augmented generation framework. Evaluated on 23 queries spanning ten intent categories, the approach achieves an F1 score of 0.632, significantly outperforming a customized processor (0.472), and further exposes a systematic bias in existing entity-level F1 metrics when assessing structured queries.

graph retrievalknowledge graphsquery understanding

This work addresses the limitations of existing retrieval-augmented generation methods, which struggle to simultaneously support structured constraints and multi-hop reasoning, as well as knowledge graph–based approaches that suffer from semantic fragmentation, high maintenance costs, and difficulty in updating. The authors propose the SAG architecture, which organizes documents via an event–entity index, preserving n-ary relations within original text blocks without constructing a global knowledge graph. At query time, SAG dynamically connects relevant event blocks through shared entities to form an evidence neighborhood. A novel dynamic hyperedge mechanism is introduced to avoid decomposing relations into triples, thereby maintaining semantic integrity while enabling efficient incremental updates and complex reasoning. By integrating SQL-style structured retrieval with dynamic connection, SAG achieves state-of-the-art performance on HotpotQA, 2WikiMultiHopQA, and MuSiQue, attaining a Recall@5 of 80.36% on MuSiQue—11.52 percentage points above the strongest baseline.

incremental updatesknowledge graphsmulti-hop reasoning

Traditional semantic search struggles to model the hierarchical structures and multi-hop cross-references prevalent in enterprise documents, limiting retrieval accuracy. This work proposes an agent-driven, recursive knowledge graph construction approach that automatically parses substitutional logic and cross-level references among documents to generate a structured graph representation. The resulting knowledge graph is integrated into a retrieval-augmented generation (RAG) framework to enable precise querying of complex regulatory logic. Evaluated on the Code of Federal Regulations benchmark, the proposed method achieves a 70% improvement in question-answering accuracy over standard vector-based RAG systems, substantially overcoming the limitations of conventional semantic retrieval.

enterprise documentsKnowledge Graphmulti-hop references

This study addresses context overflow and retrieval-generation mismatch challenges in Retrieval-Augmented Generation (RAG) systems across diverse deployment scenarios. The authors introduce an evaluation framework tailored for semi-structured knowledge bases, encompassing nine standardized application settings ranging from basic document retrieval to integrated graph-agent workflows. They propose a novel context engineering approach that combines hybrid text-graph retrieval, domain-specific knowledge graph integration, multi-step agent planning, and an agent-graph collaborative architecture. This methodology significantly reduces token consumption in GraphRAG and Agentic RAG by 19%–53% and uncovers the “retrieval-generation gap” phenomenon—demonstrating that expanded retrieval does not necessarily improve generation quality. The findings offer data-driven guidance for selecting production-grade RAG system configurations.

Agentic RAGcontext optimizationGraphRAG

Hot Scholars

JC

Jiaoyan Chen

Department of Computer Science, University of Manchester
Knowledge GraphOntologyMachine LearningLarge Language Model
JZ

Jeff Z. Pan

Professor of Knowledge Computing, University of Edinburgh
Artificial IntelligenceKnowledge Representation and ReasoningKnowledge Based Learning
AB

Andreas Both

Professor at Leipzig University of Applied Sciences & Head of Research at DATEV eG
Web EngineeringAISoftware EngineeringQuestion Answering
AS

Ashley Suh

MIT Lincoln Laboratory
Human-centered AIVisual AnalyticsKnowledge GraphsEvaluation
CL

Chengkai Li

Professor of Computer Science and Engineering, The University of Texas at Arlington
Big Data & Data ScienceComputational JournalismData-Driven Fact-CheckingNatural Language Processing