🤖 AI Summary
This study addresses the limitation of existing large language model evaluation benchmarks in assessing a model’s metacognitive awareness—particularly its ability to recognize its own errors and avoid overconfidence in localized tasks. Drawing on Flavell’s and Nelson-Narens’ theories of metacognition, the authors introduce a novel, observable confidence–accuracy alignment diagnostic framework, operationalized as a calibration assessment tool spanning five behavioral dimensions (T1–T5) and 15 measurable slots. Experiments across eight state-of-the-art models and 69 human participants demonstrate that this approach effectively uncovers calibration blind spots invisible to conventional benchmarks. For instance, while Gemini 2.5 Flash exhibits strong within-task calibration (ρ = +0.551, score: 88), it shows markedly poor cross-task difficulty prediction (score: 41), revealing a pronounced intra-model dissociation in metacognitive competence.
📝 Abstract
The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct dimensions: confidence calibration (T1-CC), epistemic vigilance (T2-EV), knowledge boundary (T3-KB), calibration range (T4-CR), and reasoning-chain validation (T5-RCV). It is evaluated on N=8 frontier models and N=69 humans. The instrument is motivated by Flavell (1979) and Nelson and Narens (1990) but operates on observable confidence-correctness alignment; it is not a validated cross-species metacognition scale, and the pre-specified human developmental hypothesis was falsified.
Composite benchmarks (MMLU, BIG-Bench, HELM, GPQA) ask whether a model produces a correct response. They are silent on whether the model knows when its response is wrong. A model can score 80 on a composite calibration benchmark and still be wildly overconfident in narrow pockets the aggregate cannot surface. The Metacognitive Probe surfaces those pockets.
Our headline is a 47-point within-model dissociation in Gemini 2.5 Flash: panel-best within-task calibration (T1-CC = 88; Spearman rho = +0.551, 95% CI [+0.14, +0.80], p = 0.005) and panel-worst cross-task difficulty prediction (T4-CR = 41; sigma_conf = 1.4 across twelve factoids).