Reasoning Models Are Accurate but Unsound on Identification

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of reasoning models to yield false affirmatives for non-identifiable queries in causal identification, alongside the absence of reliable evaluation metrics. To this end, it proposes CERTID, a formalized pipeline that leverages structural causal models and the ID algorithm to certify identifiability and verify derived expressions. Through theoretical results, the framework mitigates structural leakage, rectifies non-identifiable queries, and establishes a rigorous hierarchical guarantee mechanism. Empirical findings reveal that accuracy is not a reliable proxy for trustworthiness, with spurious claim rates varying up to 17-fold across models. Notably, CERTID achieves 97–100% decision accuracy on novel graphs, effectively resolving the longstanding challenges of evaluation deficiency and scoring difficulty in this domain.
📝 Abstract
A reasoning model asked whether a causal effect is recoverable from observational data can fail in two ways: it refuses an identifiable query or answers a nonidentifiable one. The latter is more consequential, as no observational data can validate the claimed formula. Measuring this failure requires queries that are provably non-identifiable, which prior evaluations lack, and grading that accepts correct formulas in any equivalent form, which string matching cannot provide. We build CERTID, a formal identification pipeline that addresses both limitations. CERTID uses the sound and complete causal identification algorithm ID to certify whether an effect is identifiable from a given graph and query, and verifies returned formulas against structural causal models whose interventional distributions are known exactly. CERTID further develops theoretical results to mitigate structural leakage, repair non-identifiable queries, and establish grading guarantees. We evaluate three frontier reasoning models (Gemini Flash, Gemini Pro, and GPT5.5) on 1,200 certified instances spanning 4 to 50 vertices. Accuracy proves a poor proxy for soundness: on identical instances, the false-claim rate on non-identifiable queries varies by seventeen-fold across models. We also find that models decide identifiability with 97-100% accuracy on graphs generated after the strongest model's training snapshot. Instances, the certification procedure, the verifier, and per-instance records are available at https://anonymous.4open.science/r/certid-D718.
Problem

Research questions and friction points this paper is trying to address.

causal identification
reasoning models
soundness evaluation
non-identifiable queries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Identification
Reasoning Models
Formal Verification
Structural Causal Models
CERTID
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arman Behnam
Department of Computer Science, Illinois Institute of Technology, Chicago, IL, USA
Binghui Wang
Binghui Wang
Assistant Professor, Illinois Institute of Technology
Trustworthy Machine LearningMachine LearningData Science