graph-based retrieval

Designs and implements retrieval systems that index and search nodes, paths, and subgraphs in knowledge graphs (including temporal KGs), using techniques such as graph-based indexing, random-walk or walk-based retrieval, path-finding, and temporal filtering to fetch relevant event chains or multi-hop evidence. Builds ranking and selection components that order nodes/paths/subgraphs by structural relevance or centrality and produce evidence subgraphs suitable for downstream consumers (e.g., LLM augmentation).

graph-basedretrieval

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Mixture of Structural-and-Textual Retrieval over Text-rich Graph Knowledge Bases

Feb 27, 2025
YL
Yongjia Lei
🏛️ University of Oregon | Michigan State University | Adobe | Pacific Northwest National Laboratory

Existing retrieval methods for text-rich graph knowledge bases (TG-KBs) suffer from a disconnect between structural and textual knowledge access; hybrid approaches either neglect structural retrieval or underutilize neighborhood interactions. Method: We propose the Planning–Reasoning–Organizing (PRO) framework—the first to integrate *textualized planning graphs* into retrieval. PRO employs query-driven structural traversal planning, jointly performs graph path reasoning and fine-grained text matching, and leverages structural trajectories to enhance candidate re-ranking—enabling dynamic cross-modal alignment and complementary knowledge enhancement. The method unifies interpretable planning graph generation, multi-stage neural ranking, and semantic-structural joint modeling. Contribution/Results: PRO achieves significant improvements over state-of-the-art methods across multiple TG-KB benchmarks, empirically validating both the effectiveness of structural trajectory modeling for re-ranking and the robustness and generalizability of the hybrid retrieval paradigm.

Enhances retrieval in Text-rich Graph Knowledge BasesIntegrates structural and textual knowledge retrievalProposes Planning-Reasoning-Organizing framework for queries

KG-Retriever: Efficient Knowledge Indexing for Retrieval-Augmented Large Language Models

Dec 07, 2024
WC
Weijie Chen
🏛️ Beijing University of Posts and Telecommunications | University of Science and Technology Beijing | Xiaomi Corporation

To address information fragmentation and cross-document reasoning challenges in complex retrieval tasks such as multi-hop question answering, this paper proposes HierRAG, a knowledge graph–driven hierarchical retrieval-augmented framework. HierRAG constructs a layered index graph integrating a knowledge graph layer and a collaborative document layer, leveraging graph neural networks to jointly model entity–document relationships—enabling coordinated coarse-grained semantic navigation and fine-grained knowledge localization. Unlike conventional flat RAG architectures, HierRAG introduces the first hierarchical indexing structure, significantly improving intra- and inter-document connectivity and multi-hop reasoning capability. Evaluated on five mainstream multi-hop QA benchmarks, HierRAG achieves substantial gains in both retrieval accuracy and response efficiency, demonstrating its effectiveness and generalizability in complex reasoning scenarios.

Addresses information fragmentation via hierarchical knowledge indexingEnhances multi-hop question answering in retrieval-augmented LLMsImproves cross-document retrieval efficiency using graph structures

This work addresses the semantic mismatch caused by structural heterogeneity in knowledge graph question answering and the lack of global structural awareness in existing methods. It reframes multi-hop reasoning as a schema-guided graph search task, constructing an adaptive query schema graph via a semantic-structure projection mechanism. By integrating a Triple-Dependent Graph Neural Network, the approach enables globally guided node anchoring and subgraph retrieval, thereby incorporating global structural information into the retrieval phase for the first time to generate high-quality evidence reasoning graphs. This strategy substantially improves both retrieval accuracy and evidence completeness of multi-hop reasoning paths, achieving state-of-the-art performance across multiple benchmark datasets.

Knowledge GraphMulti-hop ReasoningReasoning Path Retrieval

This work addresses the limitations of existing retrieval-augmented generation methods, which struggle to simultaneously support structured constraints and multi-hop reasoning, as well as knowledge graph–based approaches that suffer from semantic fragmentation, high maintenance costs, and difficulty in updating. The authors propose the SAG architecture, which organizes documents via an event–entity index, preserving n-ary relations within original text blocks without constructing a global knowledge graph. At query time, SAG dynamically connects relevant event blocks through shared entities to form an evidence neighborhood. A novel dynamic hyperedge mechanism is introduced to avoid decomposing relations into triples, thereby maintaining semantic integrity while enabling efficient incremental updates and complex reasoning. By integrating SQL-style structured retrieval with dynamic connection, SAG achieves state-of-the-art performance on HotpotQA, 2WikiMultiHopQA, and MuSiQue, attaining a Recall@5 of 80.36% on MuSiQue—11.52 percentage points above the strongest baseline.

incremental updatesknowledge graphsmulti-hop reasoning

To address the challenge of simultaneously achieving high retrieval efficiency, accurate reasoning, and effective hallucination suppression when augmenting large language models (LLMs) with knowledge graphs (KGs), this paper proposes SubgraphRAG—a fine-tuning-free, dynamic subgraph retrieval framework. Its core innovation lies in a lightweight MLP coupled with a parallel triple-scoring mechanism that explicitly encodes directed structural distances within the KG. Crucially, SubgraphRAG adaptively adjusts subgraph size based on query difficulty and LLM capability, enabling joint optimization of retrieval granularity and reasoning efficacy. The method integrates KG subgraph retrieval, structure-aware scoring, and LLM-coordinated generation. Evaluated on WebQSP and ComplexWebQuestions (CWQ), SubgraphRAG significantly reduces hallucination rates while improving answer accuracy and interpretability. It achieves state-of-the-art performance with Llama3.1-8B-Instruct and sets new benchmark records with GPT-4o on both datasets.

Dynamic Information RetrievalKnowledge Graph OptimizationLanguage Model Accuracy

Latest Papers

What's happening recently
View more

This work addresses the efficiency and scalability challenges posed by the irregular structure of knowledge graphs in multi-hop compositional query answering. The authors propose SeedER, a novel framework that uniquely integrates a seed-expansion mechanism with reinforcement learning. It first generates compact sets of core seed entities through lightweight dense and sparse retrieval, then iteratively expands these seeds via a graph-aware policy, decomposing global reasoning into reusable local decisions at low computational cost. By maintaining concise candidate sets while substantially improving recall, SeedER outperforms strong baselines and functions as an efficient single-stage retriever. The approach also offers theoretical advantages in compositional generalization and submodular optimization under graph constraints.

compositional queriesgraph expansionknowledge graphs

Traditional vector retrieval struggles to handle complex queries requiring structured reasoning in industrial knowledge graphs. This work constructs an aerospace supply chain knowledge graph comprising 46 node types and 64 relation types, and introduces the “operator vocabulary” hypothesis, positing that the bottleneck in graph reasoning lies not in model intelligence but in the availability of computational primitives. Guided by this insight, the authors design an LLM-driven query planner that integrates nine graph traversal primitives and six graph computation tools, forming a structure-aware retrieval-augmented generation framework. Evaluated on 23 queries spanning ten intent categories, the approach achieves an F1 score of 0.632, significantly outperforming a customized processor (0.472), and further exposes a systematic bias in existing entity-level F1 metrics when assessing structured queries.

graph retrievalknowledge graphsquery understanding

Existing knowledge graph multi-hop retrieval systems struggle to simultaneously achieve efficiency, scalability, and interpretability. This work proposes a hardware-aligned symbolic retrieval framework that reformulates multi-hop reasoning as efficient hardware operations over decomposed representations of subjects, predicates, and objects. By integrating degree-aware graph partitioning, cross-partition routing, and on-demand caching, the approach enables highly efficient retrieval at billion-edge scale. It is the first method to deliver interpretable, scalable, and hardware-efficient multi-hop retrieval on both CPUs and GPUs, significantly accelerating inference while preserving high fidelity. The framework has been successfully deployed in biomedical domains for collaborative reasoning between knowledge graphs and large language models.

hardware efficiencyinterpretabilityknowledge graph retrieval

Traditional semantic search struggles to model the hierarchical structures and multi-hop cross-references prevalent in enterprise documents, limiting retrieval accuracy. This work proposes an agent-driven, recursive knowledge graph construction approach that automatically parses substitutional logic and cross-level references among documents to generate a structured graph representation. The resulting knowledge graph is integrated into a retrieval-augmented generation (RAG) framework to enable precise querying of complex regulatory logic. Evaluated on the Code of Federal Regulations benchmark, the proposed method achieves a 70% improvement in question-answering accuracy over standard vector-based RAG systems, substantially overcoming the limitations of conventional semantic retrieval.

enterprise documentsKnowledge Graphmulti-hop references

Hot Scholars

YG

Yunjun Gao

Professor of Computer Science, Zhejiang University
DatabaseBig Data Management and Analyticsand AI Interaction with DB Technology
XS

Xing Sun

Tencent Youtu Lab
LLMMLLMAgent
YZ

Yifan Zhu

Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
QZ

Qinggang Zhang

The Hong Kong Polytechnic University
Knowledge GraphsLarge Language ModelsRetrieval-Augmented GenerationText-to-SQL
DY

Di Yin

Tencent
LLMNLPMLLM