🤖 AI Summary
In the post-training of large language models, representational anisotropy is frequently mischaracterized as a deficiency, leaving its functional specialization and interaction mechanisms with supervised fine-tuning (SFT) and reinforcement learning (RL) poorly understood. This work reveals that a small subset of channels constitutes the foundation for linguistic coherence. Accordingly, we propose SphereGate, which safeguards these core channels through a constrained gain mechanism while achieving efficient parameter-efficient fine-tuning via residual anomaly detection and activation-weighted gradient constraints. Introducing merely 0.1M additional parameters, our method surpasses baselines by 2 to 7.3 points on MATH-500, yielding performance comparable to full-model GRPO. Ultimately, this study establishes a novel paradigm for understanding and leveraging representational anisotropy in large model post-training.
📝 Abstract
LLM post-training combines supervised fine-tuning (SFT), a mode-covering forward-KL objective, with reinforcement learning (RL), a mode-seeking reverse-KL objective. Frequency-weighted likelihood training leaves a well-known signature: \emph{anisotropy}, in which a few residual channels carry disproportionately large activations. Anisotropy is widely documented and usually treated as a defect, yet its function and its interaction with post-training remain unclear. We first analyze it. A label-free outlier rule isolates about 5\% of residual channels that are essential for language modeling: removing them raises perplexity from 10 to over $10^6$, versus 35 for count-matched random channels. Yet they barely distinguish correct from incorrect reasoning. SFT reshapes them, whereas RL leaves them largely intact and adapts the complementary channels. These channels therefore form the model's \emph{coherence substrate}, and reasoning adaptation happens elsewhere. We then exploit this. \textsc{SphereGate} learns one bounded gain per residual channel on a frozen backbone. Its activation-weighted gradients provably limit movement of high-energy coherence channels and leave the remaining channels free. With 0.1M trainable parameters, \textsc{SphereGate} outperforms parameter-efficient baselines by 2.0--7.3 points on MATH-500 across Qwen2.5 (0.5B--7B) and Llama-3-8B, is comparable or exceeds full-model GRPO. Anisotropy is not a defect but a division of labor that post-training can exploit.