T3: Benchmarking Sycophancy and Skepticism in Causal Judgment

📅 2026-01-13
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates judgment biases of large language models (LLMs) in causal reasoning, revealing a “skepticism trap” at Level 1 of Pearl’s causal hierarchy—where models like Claude Haiku erroneously reject 60% of valid causal chains—and a “non-monotonic scaling paradox” at Level 3, exemplified by GPT-5.2 underperforming GPT-4-Turbo by 55 points. To address these issues, we introduce the T3 benchmark, grounded in Pearl’s ladder of causation and comprising 454 expert-crafted scenarios, assessing model performance across utility, safety, and calibrated refusal. We further propose a structured process verification protocol (RCA) and adversarial fuzzy counterfactual testing, complemented by high-resolution failure analysis, which collectively enhance both decisiveness and accuracy in causal judgments.

Technology Category

Reasoning under Uncertainty: CausalityMachine Learning: Causal LearningKnowledge Representation and Reasoning: Action, Change, and Causality

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
We introduce T3 (Testing Trustworthy Thinking), a diagnostic benchmark designed to rigorously evaluate LLM causal judgment across Pearl's Ladder of Causality. Comprising 454 expert-curated vignettes, T3 prioritizes high-resolution failure analysis, decomposing performance into Utility (sensitivity), Safety (specificity), and Wise Refusal on underdetermined cases. By applying T3 to frontier models, we diagnose two distinct pathologies: a"Skepticism Trap"at L1 (where safety-tuned models like Claude Haiku reject 60% of valid links) and a non-monotonic Scaling Paradox at L3. In the latter, the larger GPT-5.2 underperforms GPT-4-Turbo by 55 points on ambiguous counterfactuals, driven by a collapse into paralysis (excessive hedging) rather than hallucination. Finally, we use the benchmark to validate a process-verified protocol (RCA), showing that T3 successfully captures the restoration of decisive causal judgment under structured verification.
Problem

Research questions and friction points this paper is trying to address.

causal judgment
sycophancy
skepticism
large language models
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

causal judgment
benchmarking
sycophancy
skepticism
scaling paradox
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Edward Y. Chang
Computer Science, Stanford University