π€ AI Summary
This study addresses whether information in pre-trained models is erased, rerouted, or scaled during post-trainingβa question long confounded in representation compression research. We propose an identifiable information-fate framework that leverages reward null-space analysis, spectral compression techniques, and Bayesian belief-state theory to achieve causal disentanglement through controlled experiments. By introducing the concept of "causal quotient," we reveal that post-training predominantly reroutes or scales information rather than erasing it. Furthermore, we demonstrate that without KL anchoring, decision pathways are deactivated while representations persist, and contextual belief geometry remains decodable despite spectral compression in later layers. Ultimately, our findings establish that the optimization process determines the final fate of residual information.
π Abstract
What happens to information a pretrained model already encodes when post-training no longer rewards using it? The common language of representation compression conflates three fates: information may be erased, rerouted away from the decision while still represented, or rescaled to occupy less variance while still represented and used. We make these fates identifiable in models whose pretraining recovers Bayesian belief states. A reward that reads only a coarse function of the hidden state defines an exact reward-null kernel. The kernel lets us separately measure whether the information remains recoverable, whether decisions causally depend on it, and how much activation variance it occupies. Theory says what is protected: KL-anchored reinforcement learning preserves the reference policy's log-odds among equally rewarded outputs, supervised and unanchored objectives carry no such constraint, and spectral compression implies neither erasure nor loss of use. In controlled worlds, post-training mostly reroutes or rescales reward-null information and leaves it decodable. Without an anchor decisions can stop using it although the representation survives, and with one they keep using it. Erasure appears only under prolonged weight decay, for distinctions that neither reward nor next-token prediction can see. Open language models show the same dissociation: in-context belief geometry stays decodable under late-layer spectral compression, and within-class behavior depends on the anchor. Post-training thus selects a causal quotient of the pretrained belief state: the reward defines decision-equivalence, the anchor and the state update protect part of what it ignores, and optimization decides whether the rest is erased, rerouted, or rescaled.