🤖 AI Summary
This work addresses the disconnect between retrieval and generation stages in repository-level code completion, where existing methods often overlook generative potential. We propose CARET, a training-free framework that pioneers a synergistic optimization paradigm integrating retrieval and generation. Specifically, CARET introduces a sample-consistency-based cascaded retrieval routing mechanism to dynamically filter contexts, leverages KV cache reuse to reduce multi-sampling overhead, and employs reverse context likelihood for self-scoring candidate code. Evaluated across six models and multiple benchmarks, CARET improves average exact match rates by 10.8 and 5.3 percentage points over greedy decoding and self-consistency, respectively, while incurring a computational cost of only approximately 1.45× that of single-pass generation.
📝 Abstract
Retrieval-augmented generation (RAG) dominates repository-level code completion: it retrieves cross-file context (R), then decodes one greedy completion (G). Existing work mainly focuses on retrieval and stops there. We argue both stages can be improved together, with generation in particular gaining from test-time scaling. We present CARET, a training-free method. For R, CARET routes among retrieval contexts using the agreement among its own samples, cascading to an alternative context when the samples scatter. For G, it samples candidates over a cached prompt prefix, so the long retrieved context is encoded once rather than once per sample. It then selects the final completion by reverse-context likelihood: a correct completion makes the code after the cursor more probable, so the same model grades its own candidates by reading ahead. Across CrossCodeEval, RepoEval-Line, and RepoEval-API with six code models (1.1B to 7B, four families), CARET improves exact match in all 18 combinations by 10.8 points on average over greedy decoding and 5.3 over self-consistency@10. Token-level compute stays near 1.45 times one generation (measured wall-clock 1.3-2.9 times, growing with generator size). Improving retrieval and generation together yields more accurate code than improving retrieval alone, at a budget that stays close to a single pass.