🤖 AI Summary
This study addresses the tendency of reasoning models to yield false affirmatives for non-identifiable queries in causal identification, alongside the absence of reliable evaluation metrics. To this end, it proposes CERTID, a formalized pipeline that leverages structural causal models and the ID algorithm to certify identifiability and verify derived expressions. Through theoretical results, the framework mitigates structural leakage, rectifies non-identifiable queries, and establishes a rigorous hierarchical guarantee mechanism. Empirical findings reveal that accuracy is not a reliable proxy for trustworthiness, with spurious claim rates varying up to 17-fold across models. Notably, CERTID achieves 97–100% decision accuracy on novel graphs, effectively resolving the longstanding challenges of evaluation deficiency and scoring difficulty in this domain.
📝 Abstract
A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at https://anonymous.4open.science/r/certid-D718.