The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMs

📅 2026-05-10
📈 Citations: 0
Influential: 0
📄 PDF

career value

178K/year
🤖 AI Summary
This study addresses the limitation of existing large language model evaluation benchmarks in assessing a model’s metacognitive awareness—particularly its ability to recognize its own errors and avoid overconfidence in localized tasks. Drawing on Flavell’s and Nelson-Narens’ theories of metacognition, the authors introduce a novel, observable confidence–accuracy alignment diagnostic framework, operationalized as a calibration assessment tool spanning five behavioral dimensions (T1–T5) and 15 measurable slots. Experiments across eight state-of-the-art models and 69 human participants demonstrate that this approach effectively uncovers calibration blind spots invisible to conventional benchmarks. For instance, while Gemini 2.5 Flash exhibits strong within-task calibration (ρ = +0.551, score: 88), it shows markedly poor cross-task difficulty prediction (score: 41), revealing a pronounced intra-model dissociation in metacognitive competence.
📝 Abstract
The Metacognitive Probe is an exploratory five-task, 15-slot diagnostic that decomposes an LLM's confidence behaviour into five behaviourally-distinct dimensions: confidence calibration (T1-CC), epistemic vigilance (T2-EV), knowledge boundary (T3-KB), calibration range (T4-CR), and reasoning-chain validation (T5-RCV). It is evaluated on N=8 frontier models and N=69 humans. The instrument is motivated by Flavell (1979) and Nelson and Narens (1990) but operates on observable confidence-correctness alignment; it is not a validated cross-species metacognition scale, and the pre-specified human developmental hypothesis was falsified. Composite benchmarks (MMLU, BIG-Bench, HELM, GPQA) ask whether a model produces a correct response. They are silent on whether the model knows when its response is wrong. A model can score 80 on a composite calibration benchmark and still be wildly overconfident in narrow pockets the aggregate cannot surface. The Metacognitive Probe surfaces those pockets. Our headline is a 47-point within-model dissociation in Gemini 2.5 Flash: panel-best within-task calibration (T1-CC = 88; Spearman rho = +0.551, 95% CI [+0.14, +0.80], p = 0.005) and panel-worst cross-task difficulty prediction (T4-CR = 41; sigma_conf = 1.4 across twelve factoids).
Problem

Research questions and friction points this paper is trying to address.

confidence calibration
metacognition
large language models
overconfidence
behavioral diagnostics
Innovation

Methods, ideas, or system contributions that make the work stand out.

metacognitive probe
confidence calibration
epistemic vigilance
reasoning-chain validation
behavioral diagnostics