🤖 AI Summary
Although large language models achieve high accuracy, their confidence calibration remains severely decoupled from error recognition. This study systematically evaluates the metacognitive monitoring capabilities of frontier models by investigating the alignment between reported confidence and actual performance through cross-model assessment, confidence ranking analysis, reference model comparison, and peer supervision mechanisms. The findings reveal a fundamental decoupling between problem-solving and self-evaluation: models exhibit near-random discrimination on challenging instances, maintain high confidence in shared errors, and derive only marginal improvements from simple prompting-based review strategies. By exposing these intrinsic limitations of self-verification mechanisms, this work provides critical insights for enhancing AI reliability and underscores the necessity of developing more robust approaches to align model confidence with factual correctness.
📝 Abstract
Reliable decisions depend on recognizing when an answer may be wrong. In biological cognition, metacognitive monitoring can dissociate from task performance, raising the question of how closely solving and judging are linked in language models. Here we study the confidence reports of four frontier models across 15 benchmarks. High task accuracy can coexist with weak error discrimination: a model solves 97% of competition mathematics problems while its answer-time confidence ranks correct answers above errors barely better than chance. Confidence separates correct answers from errors more effectively on questions solved by a separate reference model, while review brings limited improvement on reference-hard questions. Aggregate discrimination also rewards ranking correct answers on easy questions above errors on hard ones, which question-only forecasts already do well. Cross-evaluation helps most where the evaluator answered correctly, and errors shared by the two models usually retain high confidence. Hard questions and shared errors remain difficult targets for prompted self-review and peer oversight, even in models with strong problem-solving performance.