🤖 AI Summary
This work addresses the challenge of efficiently extracting sparse causal abstractions from pretrained neural networks that maintain high interventional fidelity, without resorting to brute-force interventions or retraining. By framing structured pruning as a search for approximate causal abstractions, the authors model the network as a deterministic structural causal model and introduce a closed-form pruning criterion derived from a second-order Taylor expansion of interventional risk. Under a unified curvature assumption, this criterion reduces to the activation variance method, while also clarifying the conditions under which that heuristic fails. Empirical validation via interventional swapping demonstrates that the extracted abstractions exhibit strong causal faithfulness, enabling efficient and interpretable discovery of causal structure.
📝 Abstract
Neural networks are hypothesized to implement interpretable causal mechanisms, yet verifying this requires finding a causal abstraction -- a simpler, high-level Structural Causal Model (SCM) faithful to the network under interventions. Discovering such abstractions is hard: it typically demands brute-force interchange interventions or retraining. We reframe the problem by viewing structured pruning as a search over approximate abstractions. Treating a trained network as a deterministic SCM, we derive an Interventional Risk objective whose second-order expansion yields closed-form criteria for replacing units with constants or folding them into neighbors. Under uniform curvature, our score reduces to activation variance, recovering variance-based pruning as a special case while clarifying when it fails. The resulting procedure efficiently extracts sparse, intervention-faithful abstractions from pretrained networks, which we validate via interchange interventions.