🤖 AI Summary
This study systematically evaluates judgment biases of large language models (LLMs) in causal reasoning, revealing a “skepticism trap” at Level 1 of Pearl’s causal hierarchy—where models like Claude Haiku erroneously reject 60% of valid causal chains—and a “non-monotonic scaling paradox” at Level 3, exemplified by GPT-5.2 underperforming GPT-4-Turbo by 55 points. To address these issues, we introduce the T3 benchmark, grounded in Pearl’s ladder of causation and comprising 454 expert-crafted scenarios, assessing model performance across utility, safety, and calibrated refusal. We further propose a structured process verification protocol (RCA) and adversarial fuzzy counterfactual testing, complemented by high-resolution failure analysis, which collectively enhance both decisiveness and accuracy in causal judgments.
📝 Abstract
We introduce T3 (Testing Trustworthy Thinking), a diagnostic benchmark designed to rigorously evaluate LLM causal judgment across Pearl's Ladder of Causality. Comprising 454 expert-curated vignettes, T3 prioritizes high-resolution failure analysis, decomposing performance into Utility (sensitivity), Safety (specificity), and Wise Refusal on underdetermined cases. By applying T3 to frontier models, we diagnose two distinct pathologies: a"Skepticism Trap"at L1 (where safety-tuned models like Claude Haiku reject 60% of valid links) and a non-monotonic Scaling Paradox at L3. In the latter, the larger GPT-5.2 underperforms GPT-4-Turbo by 55 points on ambiguous counterfactuals, driven by a collapse into paralysis (excessive hedging) rather than hallucination. Finally, we use the benchmark to validate a process-verified protocol (RCA), showing that T3 successfully captures the restoration of decisive causal judgment under structured verification.