🤖 AI Summary
Observationally equivalent causal models may diverge under intervention, yet existing methods lack consistency guarantees. This work proposes the Interventional Separation Selection (ISS) algorithm, which iteratively queries the true system to eliminate contradictory candidate models. ISS employs mixed-integer linear programming to handle the infinite version space of continuous variables and introduces a verifiable stopping condition that requires no ground-truth labels. Consequently, it certifies interventional consistency between surviving models and the true system within finite cost. Experiments on MNIST causal abstraction demonstrate that ISS achieves certification with an average of only 13.6 interventions. Furthermore, we find that certificates fail when networks bypass critical units, and random interventions successfully refute 69% of such spurious certificates.
📝 Abstract
Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the true system with an admissible intervention on which the surviving candidate models disagree, discards the candidates the outcome contradicts, and stops once no intervention within a cost bound separates the survivors. If the true system is among the candidates, this stopping condition certifies that every survivor agrees with it on every admissible intervention within the bound, a guarantee that no observational learner can give, however much data it sees. The stopping condition depends only on the survivors, so it can be checked without knowing the truth. For continuous variables the candidates form an infinite version space, and mixed-integer linear programs decide the stopping condition exactly over all of it, with agreement holding up to a tolerance. On a three-digit colored MNIST causal abstraction task in which ink hue tracks digit size, plain convolutional networks trained on examples reach zero held-out error, yet disagree with shape-based labels on 26% of single-digit edits, as often as hue-based labels do. Auditing the causal abstractions of networks observed only on such images, ISS certifies what each network perceives with 13.6 interventions per image on average, and each certificate, checked against every admissible intervention, holds whenever the network's true abstraction is among the candidates. When a network bypasses a unit that every candidate abstraction relies on, certificates covering interventions on that unit can be silently void, and twenty random validation interventions refute 69% of them.