🤖 AI Summary
This study addresses the challenge faced by large code models in selectively forgetting specific implementations without compromising general programming capabilities. To this end, we propose UNBIND, a framework that enables inference-time selective unlearning without modifying model weights. By decoupling the construction of hidden-state targets from suppression directions, UNBIND achieves precise and efficient unlearning through inference-time directional steering under fixed parameters. Experimental results demonstrate that UNBIND reduces the reproduction rate of target code by over 97%, significantly outperforming existing baselines, while maximally preserving general-purpose performance on standard benchmarks such as HumanEval.
📝 Abstract
Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patterns, creating a tension between forgetting specific implementations and preserving general programming ability. We propose \textbf{UNBIND}, a code unlearning framework that separately considers which hidden states correspond to the target code and how to suppress its reproduction. By constructing separate directions for these objectives, UNBIND achieves selective unlearning at inference time while keeping model weights fixed. Our evaluation covers fourteen baselines across two code models and two corpora. UNBIND achieves the highest joint forgetting and utility score in every setting. It reduces target code reproduction by 97.3\% to 99.1\% as measured by F-BLEU, with at most two fewer HumanEval+ and six fewer MBPP+ problems solved than the original models. In repeated extraction tests under a fixed budget, the number of targets yielding exact spans of at least 50 tokens falls from 188--262 to 0--2 out of 300 per setting. No extracted span reaches 100 tokens, and the mean best recovery ratio ranges from 0.43\% to 6.45\%. Multilingual and related-code evaluations further show effective forgetting with limited impact on useful programming capabilities, supporting UNBIND as a practical approach to selective code unlearning.