🤖 AI Summary
This study addresses the prohibitive computational overhead of full-data fine-tuning and the negative transfer induced by low-quality samples in code retrieval models. To this end, we propose MAP4CS, an adaptive data pruning framework that integrates syntactic structures, semantic representations, and distributional features for multi-dimensional perceptual analysis. Coupled with a rule-based filtering pipeline, it automatically adapts to corpus characteristics to perform deduplication and denoising, precisely extracting a high-quality core subset to construct an information-dense training set. Experiments demonstrate that MAP4CS embodies the "less is more" paradigm: utilizing merely 5% of the training data, it surpasses random sampling baselines and achieves performance comparable to or exceeding full-data fine-tuning, thereby significantly enhancing both the training efficiency and generalization capability of code retrieval models.
📝 Abstract
Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness.
To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.