GraphRAFT: Retrieval Augmented Fine-Tuning for Knowledge Graphs on Graph Databases

๐Ÿ“… 2025-04-07
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing GraphRAG approaches for knowledge graph (KG) question answering over graph databases commonly neglect or underutilize the retrieval step, leading to inaccurate Cypher query generation and frequent hallucinations. To address this, we propose the first plug-and-play end-to-end framework that deeply integrates retrieval-augmented generation (RAG) with large language model (LLM) fine-tuningโ€”enabling precise multi-hop Cypher query generation and verifiable reasoning. Our method unifies subgraph context construction, native graph database interfacing, LLM fine-tuning, and retrieval-augmented inference. Evaluated on two major text-attribute KG QA benchmarks, our approach consistently outperforms state-of-the-art methods across all four metrics. It further achieves high sample efficiency during training and strong system scalability, making it both practically deployable and theoretically grounded.

Technology Category

Data Mining & Knowledge Management: Linked Open Data, Knowledge Graphs & KB CompletionNatural Language Processing: Question AnsweringKnowledge Representation and Reasoning: Knowledge Acquisition

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
๐Ÿ“ Abstract
Large language models have shown remarkable language processing and reasoning ability but are prone to hallucinate when asked about private data. Retrieval-augmented generation (RAG) retrieves relevant data that fit into an LLM's context window and prompts the LLM for an answer. GraphRAG extends this approach to structured Knowledge Graphs (KGs) and questions regarding entities multiple hops away. The majority of recent GraphRAG methods either overlook the retrieval step or have ad hoc retrieval processes that are abstract or inefficient. This prevents them from being adopted when the KGs are stored in graph databases supporting graph query languages. In this work, we present GraphRAFT, a retrieve-and-reason framework that finetunes LLMs to generate provably correct Cypher queries to retrieve high-quality subgraph contexts and produce accurate answers. Our method is the first such solution that can be taken off-the-shelf and used on KGs stored in native graph DBs. Benchmarks suggest that our method is sample-efficient and scales with the availability of training data. Our method achieves significantly better results than all state-of-the-art models across all four standard metrics on two challenging Q&As on large text-attributed KGs.
Problem

Research questions and friction points this paper is trying to address.

Improves retrieval for multi-hop KG queries
Generates correct Cypher queries for subgraphs
Enhances accuracy in graph database QA
Innovation

Methods, ideas, or system contributions that make the work stand out.

Finetunes LLMs to generate Cypher queries
Retrieves high-quality subgraph contexts efficiently
Works with KGs in native graph databases
Neo4j
A
Alfred Clemedtson
Neo4j
B
Borun Shi
Neo4j