Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that language models often struggle to distinguish evidence types and perform auditable, editable causal reasoning when answering interventional questions. The authors propose a mechanism library approach with type supervision that, under a frozen fine-tuning protocol, induces discrete mechanism slots segregated by evidence type, achieving for the first time a functional disentanglement between routing and answer readout. The resulting architecture incurs negligible performance cost (quality loss of only 0.0082 nats), is exactly invertible, and fully auditable. Evaluated on the CausalWorld benchmark, the method demonstrates cross-scale efficacy (22.6M/125M parameters), with clear separation between routing and readout (output discrepancy ≤3.4×10⁻⁶), flawless execution across 250 single edits and 1,000 stacked rollbacks, and rigorous validation via preregistered, machine-verifiable criteria.
📝 Abstract
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision organizes routing, yet remains functionally decoupled from answer readout. We establish this with a typed mechanism library -- discrete mechanism slots partitioned by evidence type, auditable at the state level -- on a causal-world benchmark with exact interventional ground truth, under a frozen protocol, at two scales (22.6M and 125M). Four preregistered findings. (i) Origin. Slot-by-type organization is induced by type-level supervision: absent in architecturally identical unsupervised controls, not buyable by content-free gating labels, and statistically attributable to the supervision signal, replicating at 125M under a powered preregistered protocol (all nine cells passed). (ii) Boundary. The induced structure is a typed routing index with a sharp routing/readout boundary: slot codes scaffold routing but do not drive answer readout ($|Δ\hat{y}| \le 3.4\times10^{-6}$, zero collateral, three seeds, stable across a 5.6x scale window) -- we therefore make no behavioral-editability claim. (iii) Cost. The structure is free: LM quality matches a parameter-matched monolith within 0.0082 nats. (iv) Trust. The library state is exactly local under edit and bit-exactly revertible -- 250 single-edit and 1,000 stacked reverts per seed, zero failures. We further find that the unsupervised null itself moves with scale, so comparisons reusing a null calibrated at one scale may be confounded at another. Every claim is tied to a preregistered, machine-checkable criterion archived before the data it governs; the full audit trail, including one criterion we failed and how the frozen protocol handled it, is released as an appendix.
Problem

Research questions and friction points this paper is trying to address.

causal structure
interventional reasoning
mechanism library
routing/readout boundary
type-level supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

typed mechanism library
causal structure induction
routing/readout decoupling
functional modularity
pre-registered auditing