MAP4CS: A Multi-dimensional Data Pruning Framework for Efficient Code Retriever Fine-tuning

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive computational overhead of full-data fine-tuning and the negative transfer induced by low-quality samples in code retrieval models. To this end, we propose MAP4CS, an adaptive data pruning framework that integrates syntactic structures, semantic representations, and distributional features for multi-dimensional perceptual analysis. Coupled with a rule-based filtering pipeline, it automatically adapts to corpus characteristics to perform deduplication and denoising, precisely extracting a high-quality core subset to construct an information-dense training set. Experiments demonstrate that MAP4CS embodies the "less is more" paradigm: utilizing merely 5% of the training data, it surpasses random sampling baselines and achieves performance comparable to or exceeding full-data fine-tuning, thereby significantly enhancing both the training efficiency and generalization capability of code retrieval models.
📝 Abstract
Retrieval-Augmented Generation (RAG) has become a cornerstone in software engineering for enhancing Large Language Models (LLMs) with domain-specific knowledge. However, adapting retrievers to evolving code repositories remains challenging due to the noise and redundancy inherent in massive code corpora. Standard fine-tuning on the full corpus is computationally expensive and often leads to sub-optimal performance due to negative transfer from low-quality samples. Conversely, simple random sampling fails to guarantee data representativeness. To address these challenges, we propose MAP4CS (Multi-dimensional Awareness Pruning for Code Search), an adaptive data pruning framework. MAP4CS identifies a small, high-quality core subset by integrating syntactic structure, semantic diversity, and distributional representation, followed by a rigorous rule-based filtering pipeline. Extensive experiments on two large-scale datasets demonstrate that MAP4CS consistently outperforms random sampling baselines using only 5% of the training data. Remarkably, it achieves performance comparable to, or even superior to, fine-tuning on the full dataset, validating the ''less is more'' hypothesis in data-centric AI. Furthermore, linguistic analysis reveals an adaptive optimization mechanism: MAP4CS automatically functions as a de-duplicator for redundant corpora and a denoiser for chaotic ones, constructing a training corpus that is both lexically diverse and information-dense.
Problem

Research questions and friction points this paper is trying to address.

Retrieval-Augmented Generation
Code Retriever Fine-tuning
Data Pruning
Code Corpora
Negative Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Pruning
Code Retrieval
Retrieval-Augmented Generation
Multi-dimensional Analysis
Adaptive Optimization
🔎 Similar Papers
No similar papers found.
Y
Yuxuan Chen
School of Software Engineering, Sun Yat-sen University, China
Mingwei Liu
Mingwei Liu
Rutgers University
China laborhigh performance work systems
G
Guangsheng Ou
School of Software Engineering, Sun Yat-sen University, China
Z
Zekai Zhang
School of Software Engineering, Sun Yat-sen University, China
Z
Zike Li
School of Software Engineering, Sun Yat-sen University, China
Yanlin Wang
Yanlin Wang
Sun Yat-sen University
software engineeringNLPprogramming languages
P
Pelin Zheng
School of Software Engineering, Sun Yat-sen University, China