Evaluating Semantic and Quality-Aware Retrieval for Source Code Repositories

📅 2026-07-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional keyword-based code retrieval struggles to meet the demands of natural language queries, intent understanding, and code quality assessment. This work proposes a hybrid retrieval system that integrates semantic search with large language model (LLM)-generated quality metadata, supporting four query modes: semantic, quality-filtered, hybrid, and automatic routing. The approach innovatively incorporates function-level code slicing, text-code embeddings, and ChromaDB vector storage, and—novelly—leverages LLM-generated quality scores for dynamic query routing. Experiments on a C-language educational code corpus demonstrate strong performance: semantic retrieval achieves nDCG@5 of 0.820 and Success@5 of 0.800; automatic routing attains 100% accuracy; and in 9 out of 12 cases, LLM-predicted quality scores deviate by no more than one point from human evaluations.
📝 Abstract
Keyword-based retrieval is limited for source-code repositories when queries are expressed in natural language or concern implementation intent and code quality rather than exact tokens. This study evaluates a prototype retrieval system that combines function-level fragmentation, text-and-code embeddings, ChromaDB vector storage, LLM-derived quality metadata, and four retrieval modes: semantic, quality-filtered, hybrid, and automatic routing. The concrete evaluation uses an educational C-code corpus. The full corpus contains 563 anonymized programmer identifiers and 8,951 C files; a reproducible 10% indexed sample contains 56 programmer identifiers, 847 files, and 3,839 fragments. Across 15 manually judged queries, semantic retrieval achieved nDCG@5 of 0.820, Success@5 of 0.800, and MRR of 0.644. The automatic router selected the expected mode for all 15 queries. In a small manual audit, LLM-derived quality scores were within one point of the manual assessment for 9 of 12 fragments. Within the reported query set, semantic retrieval was the strongest overall mode, while explicit quality metadata was most useful for explicitly quality-oriented queries.
Problem

Research questions and friction points this paper is trying to address.

source code retrieval
semantic search
code quality
natural language queries
implementation intent
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic retrieval
code quality metadata
function-level fragmentation
LLM-derived embeddings
automatic routing
M
Marek Horváth
Technical University of Košice, Faculty of Electrical Engineering and Informatics, Department of Computers and Informatics, Košice, Slovakia
E
Emília Pietriková
Technical University of Košice, Faculty of Electrical Engineering and Informatics, Department of Computers and Informatics, Košice, Slovakia