CARET: Training-Free Test-Time Scaling for Repository-Level Code Completion
This work addresses the disconnect between retrieval and generation stages in repository-level code completion, where existing methods often overlook generative potential. We propose CARET, a training-free framework that pioneers a synergistic optimization paradigm integrating retrieval and generation. Specifically, CARET introduces a sample-consistency-based cascaded retrieval routing mechanism to dynamically filter contexts, leverages KV cache reuse to reduce multi-sampling overhead, and employs reverse context likelihood for self-scoring candidate code. Evaluated across six models and multiple benchmarks, CARET improves average exact match rates by 10.8 and 5.3 percentage points over greedy decoding and self-consistency, respectively, while incurring a computational cost of only approximately 1.45× that of single-pass generation.