🤖 AI Summary
This study addresses the disconnect between agent task performance and the preservation of causal mechanisms by proposing a theory of causal retention and the Causal Core framework. Through interface factorization and selective adaptation, it achieves evidence-gated writing, readout filtering, and local diagnostic updates, ensuring frozen states respond accurately to interventional probes independently of training. Theoretically, we establish a lower bound on Bayesian decision risk, prove a posterior coverage theorem and exact editing decomposition, and demonstrate that causal retention is independent of task sufficiency. Empirically, the method attains 1.000 accuracy on source mechanisms while accepting only 5.6% of candidates. Furthermore, in TD-MPC2, it improves effect-sign accuracy from 0.057 to 0.948 without compromising stability.
📝 Abstract
Task performance need not determine which intervention mechanism an agent retains. We study causal retention: whether a frozen learned state answers a mechanism-probe map fixed independently of training, including action, context, direct target, value, and delay. For finite structural causal model classes, the optimal probe error is a Bayes decision risk. It vanishes exactly when every learning-interface fiber lies within one probe-answer fiber; any state obtained by post-processing that interface inherits the same lower bound. A posterior-coverage theorem characterizes budgeted retesting, while an exact edit decomposition shows that the shifted set is the unique support of an error-free target update. Causal Core implements these conditions through evidence-gated writing, readout filtering, temporal credit, hidden-context setup, and local diagnostic updates. Experiments cover finite causal systems, continuous simulators, an official TD-MPC2 world model, and Qwen2.5-7B-Instruct. A frozen Qwen last-layer probe reaches 0.958 balanced accuracy on source mechanisms but 0.583 on changed delays; the gated mechanism state reaches 1.000 and accepts only 0.056 of synchronized-readout candidates. In TD-MPC2, five target states per actuator recover effect-sign accuracy from 0.057 to 0.948 without degrading stable responses. Causal retention is therefore distinct from task sufficiency and source-domain decodability.