code snippet extraction

Designs and builds systems that automatically locate and extract contiguous code fragments from text, documentation, or repositories, producing formatted snippets labeled with programming language and functionality metadata; outputs are suitable for retrieval, display, or execution (including multilingual or runnable variants).

codesnippetextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Current large code models exhibit limited performance in repository-level code generation due to their neglect of cross-file dependencies and structural context, while conventional NLP-based retrieval-augmented approaches struggle to effectively model the inherent structure of code. To address this, this work proposes Hydra, a novel framework that treats code as structured entities and introduces a hierarchical code tree index, a dependency-aware retriever (DAR), and a hybrid retrieval mechanism that integrates functional dependencies with semantic similarity. Hydra departs from traditional NLP-style paradigms for code processing and achieves state-of-the-art results on the DevEval and RepoExec benchmarks, surpassing the strongest baseline by over 5% in Pass@1. Notably, it enables smaller models equipped with Hydra to match the performance of larger models using conventional retrievers.

code coherencecode structurecross-file dependencies

AutoFL: A Tool for Automatic Multi-granular Labelling of Software Repositories

Aug 05, 2024
CS
Cezar Sas
🏛️ University of Groningen

Software developers face inefficient and time-consuming challenges in comprehending large, multifunctional codebases; existing README-based, coarse-grained project-level categorization fails to support fine-grained functional understanding. To address this, we propose AutoFL—the first automated, cross-granularity functional domain labeling method supporting file-, package-, and project-level annotations without relying on non-code documentation (e.g., READMEs). AutoFL directly models source code semantics via a weakly supervised learning framework that integrates code text parsing, multi-granularity semantic embedding, and hierarchical aggregation for end-to-end label generation. Evaluated across multilingual open-source projects, AutoFL significantly improves the accuracy, consistency, and interpretability of functional labels compared to baselines. It effectively alleviates key bottlenecks in software comprehension by enabling precise, scalable, and documentation-agnostic functional awareness.

Automate labeling software repositories for faster code comprehensionEnable multi-granular annotations at file, package, and project levelsImprove domain-specific file classification using source code input

CodeRAG-Bench: Can Retrieval Augment Code Generation?

Jun 20, 2024
ZZ
Z. Z. Wang
🏛️ Carnegie Mellon University | University of Washington | University of Southern California

This work systematically investigates the effectiveness and limitations of Retrieval-Augmented Generation (RAG) for code generation. Addressing the core question—*when and why does retrieval improve code generation?*—the authors introduce CodeRAG-Bench, the first large-scale, multi-scenario RAG benchmark for code, covering foundational programming, open-domain, and repository-level tasks, and integrating five heterogeneous context sources: competitive programming solutions, tutorials, documentation, Stack Overflow posts, and GitHub repositories. Through comprehensive evaluation across multiple LLMs (e.g., CodeLlama, DeepSeek-Coder) and retrievers (e.g., BM25, DPR, ColBERT), the study reveals that high-quality retrieval substantially improves generation accuracy; however, current approaches suffer from critical bottlenecks—including retrieval failure under low lexical overlap and generators’ inability to effectively incorporate short, sparse retrieved contexts. CodeRAG-Bench is publicly released to serve as a community-standard evaluation platform for code-oriented RAG research.

Assessing retrieval benefits in code generation scenarios.Exploring retrieval-augmented generation for code tasks.Identifying challenges in context retrieval for code models.

Context-Augmented Code Generation Using Programming Knowledge Graphs

Oct 09, 2024
IS
Iman Saberi
🏛️ The University of British Columbia

To address imprecise retrieval and frequent hallucinations in large language models (LLMs) and code-LLMs—caused by inadequate semantic understanding and limited context capacity in complex programming tasks—this paper proposes PKG-RAG, a programming knowledge graph (PKG)-driven fine-grained retrieval-augmented generation framework. Methodologically, it constructs a semantically enriched PKG enabling block-level and function-level retrieval; designs a tree-pruning algorithm to enhance retrieval precision; introduces a non-RAG re-ranking mechanism to suppress hallucinations; and integrates a Fill-in-the-Middle (FIM)-aware module for automated comment and docstring generation. Contributions include: (i) the first PKG-driven dual-granularity retrieval paradigm; (ii) a synergistic optimization strategy combining tree pruning and re-ranking; and (iii) FIM-aware code completion without additional training. Experiments show up to 20% absolute improvement in pass@1 on HumanEval and a 34% gain over SOTA on MBPP, significantly enhancing robustness on complex tasks and reducing hallucination rates.

Enhances retrieval precision with tree-pruning and re-rankingImproves code generation by using Programming Knowledge GraphsReduces hallucinations in retrieval-augmented generation models

Latest Papers

What's happening recently
View more

This study addresses the lack of systematic guidance for component-level configuration in retrieval-augmented generation (RAG) systems applied to software engineering tasks, a gap that has led practitioners to rely on costly trial-and-error approaches. Through a large-scale empirical evaluation, the work systematically assesses the impact of four core RAG modules—query processing, retrieval models, context refinement, and generators—across code generation, summarization, and repair tasks. The experiments span four query processing techniques, seven retrieval models, four refinement methods, and six generators, yielding 21 distinct configurations. Notably, the study reveals that retrieval components—particularly the choice of retrieval algorithm—often exert a greater influence on overall performance than the generator itself, with BM25 demonstrating consistently robust results across multiple tasks. These findings provide a data-driven prioritization framework for optimizing RAG systems in software engineering, substantially reducing configuration costs.

Component ConfigurationEmpirical StudyRAG Optimization

This work addresses the limitations of large language models (LLMs) in warehouse-scale code completion, where cross-file dependencies and constrained context windows hinder performance, while existing retrieval-augmented approaches based on semantic indexing or graph structures incur high computational overhead. The paper presents the first systematic investigation into lightweight, index-free lexical retrieval—specifically using tools like ripgrep—for this task, introducing GrepRAG. The method first employs an LLM to automatically generate ripgrep queries (Naive GrepRAG), then enhances retrieval quality through identifier-weighted ranking and a structure-aware deduplication mechanism. Evaluated on CrossCodeEval and RepoEval-Updated, GrepRAG significantly outperforms current state-of-the-art methods, achieving relative improvements of 7.04%–15.58% in exact code match accuracy, thereby demonstrating the effectiveness of efficient lexical retrieval for code completion.

code completioncross-file dependenciesindex-free retrieval

Traditional keyword-based code retrieval struggles to meet the demands of natural language queries, intent understanding, and code quality assessment. This work proposes a hybrid retrieval system that integrates semantic search with large language model (LLM)-generated quality metadata, supporting four query modes: semantic, quality-filtered, hybrid, and automatic routing. The approach innovatively incorporates function-level code slicing, text-code embeddings, and ChromaDB vector storage, and—novelly—leverages LLM-generated quality scores for dynamic query routing. Experiments on a C-language educational code corpus demonstrate strong performance: semantic retrieval achieves nDCG@5 of 0.820 and Success@5 of 0.800; automatic routing attains 100% accuracy; and in 9 out of 12 cases, LLM-predicted quality scores deviate by no more than one point from human evaluations.

code qualityimplementation intentnatural language queries

This work addresses the limited effectiveness of general-purpose large language models in code completion tasks within private enterprise codebases, where domain-specific structures and coding styles hinder performance. To overcome this challenge, the authors propose a semantic-scope-based approach for automatically constructing training data, which integrates retrieval-augmented generation (RAG) with supervised fine-tuning to efficiently customize medium-scale language models. Experimental results on two real-world enterprise codebases demonstrate that the resulting customized models significantly outperform larger, unadapted general-purpose models in code completion accuracy, while maintaining strong generalization capabilities on public benchmarks. These findings validate both the efficacy and practicality of the proposed method for adapting language models to proprietary software environments.

code completionlarge language modelsmodel customization

Hot Scholars

MC

Minghao Chen

Hangzhou Dianzi University
Deep LearningDomain AdaptationVision and LanguageLLM Agents
JJ

Junhao Jia

Hangzhou Dianzi University
Explainable AI (XAI)Interpretable Computer VisionMedical Image Analysis
LC

Lulu Chen

Virginia Tech
Machine LearningData MiningBioinformatics
AK

Ankur Kumar

University of California Los Angeles
MB

Murat Bilgehan Ertan

CWI, Vrije Universiteit Amsterdam
machine learningcomputer securityprivacy