🤖 AI Summary
This study addresses the unintended behavioral drift in language models caused by task-specific fine-tuning by proposing the ATLAS framework. This method pioneers the transformation of the geometric structure of preserved-domain activation maps into local reference centers and directional filters, constructing input-dependent adaptation rules within a shared low-rank residual space. Coupled with KL-divergence optimization, it dynamically constrains model updates during both training and inference, thereby harmonizing skill acquisition with behavioral preservation. Experiments demonstrate that ATLAS significantly reduces retained-output deviation across multiple backbones, including Qwen3-8B, outperforming seven baseline methods while maintaining compact storage requirements and minimal decoding overhead.
📝 Abstract
Task-specific fine-tuning can rewrite a language model's answers beyond the training task, complicating updates that must preserve existing behavior. We introduce ATLAS, which turns retained-domain representations into an input-dependent rule for task adaptation. An activation atlas supplies local reference centers and directional filters to a shared low-rank residual. Target supervision learns the residual, while retained geometry shapes its action throughout training and inference. On Qwen3-8B, ATLAS achieves lower mean retained-output Kullback-Leibler (KL) divergence than all seven published baselines at shared coding-performance requirements, with consistent advantages across multiple training seeds. Structural comparisons identify the contributions of retained reference states and directional conditioning, and answer-level analyses show fewer rewritten mathematical answers and more stable commonsense choices. Experiments spanning five backbones and two retained domains further demonstrate coding gains with reduced retained-output movement. With compact storage and modest decoding overhead, ATLAS provides a practical mechanism for acquiring specialized skills while maintaining continuity in existing responses.