Context-Augmented Code Generation Using Programming Knowledge Graphs

📅 2024-10-09
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address imprecise retrieval and frequent hallucinations in large language models (LLMs) and code-LLMs—caused by inadequate semantic understanding and limited context capacity in complex programming tasks—this paper proposes PKG-RAG, a programming knowledge graph (PKG)-driven fine-grained retrieval-augmented generation framework. Methodologically, it constructs a semantically enriched PKG enabling block-level and function-level retrieval; designs a tree-pruning algorithm to enhance retrieval precision; introduces a non-RAG re-ranking mechanism to suppress hallucinations; and integrates a Fill-in-the-Middle (FIM)-aware module for automated comment and docstring generation. Contributions include: (i) the first PKG-driven dual-granularity retrieval paradigm; (ii) a synergistic optimization strategy combining tree pruning and re-ranking; and (iii) FIM-aware code completion without additional training. Experiments show up to 20% absolute improvement in pass@1 on HumanEval and a 34% gain over SOTA on MBPP, significantly enhancing robustness on complex tasks and reducing hallucination rates.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Large Multimodal Models (LMMs)Data Mining & Knowledge Management: Conversational Systems for Recommendation & Retrieval

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Large Language Models (LLMs) and Code-LLMs (CLLMs) have significantly improved code generation, but, they frequently face difficulties when dealing with challenging and complex problems. Retrieval-Augmented Generation (RAG) addresses this issue by retrieving and integrating external knowledge at the inference time. However, retrieval models often fail to find most relevant context, and generation models, with limited context capacity, can hallucinate when given irrelevant data. We present a novel framework that leverages a Programming Knowledge Graph (PKG) to semantically represent and retrieve code. This approach enables fine-grained code retrieval by focusing on the most relevant segments while reducing irrelevant context through a tree-pruning technique. PKG is coupled with a re-ranking mechanism to reduce even more hallucinations by selectively integrating non-RAG solutions. We propose two retrieval approaches-block-wise and function-wise-based on the PKG, optimizing context granularity. Evaluations on the HumanEval and MBPP benchmarks show our method improves pass@1 accuracy by up to 20%, and outperforms state-of-the-art models by up to 34% on MBPP. Our contributions include PKG-based retrieval, tree pruning to enhance retrieval precision, a re-ranking method for robust solution selection and a Fill-in-the-Middle (FIM) enhancer module for automatic code augmentation with relevant comments and docstrings.
Problem

Research questions and friction points this paper is trying to address.

Improves code generation by using Programming Knowledge Graphs
Reduces hallucinations in retrieval-augmented generation models
Enhances retrieval precision with tree-pruning and re-ranking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Leverages Programming Knowledge Graph for semantic retrieval
Uses tree-pruning to reduce irrelevant context
Integrates re-ranking to minimize hallucinations
I
Iman Saberi
Department of Computer Science, Mathematics, Physics and Statistics, The University of British Columbia, Kelowna, BC, Canada
F
Fatemeh Fard
Department of Computer Science, Mathematics, Physics and Statistics, The University of British Columbia, Kelowna, BC, Canada