Diagnosing Correctness Probes under Self-Judgement Confounding

πŸ“… 2026-07-18
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing correctness probes struggle to disentangle objective correctness (OC) from the model’s self-judgment (SJ) in language model hidden states due to their high empirical alignment. This work constructs samples where OC and SJ diverge and employs a factorization-based approach to isolate and estimate the directions associated with each signal, evaluating their polarity and transferability across mathematical reasoning and factual recall tasks. Experiments on four instruction-tuned models (up to 14B parameters) reveal that current probes predominantly capture SJ rather than OC; furthermore, the SJ-associated direction exhibits strong cross-task transferability, whereas the OC direction performs worse than random. These findings systematically demonstrate, for the first time, a fundamental limitation in the semantic reliability of prevailing correctness probes.
πŸ“ Abstract
Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.
Problem

Research questions and friction points this paper is trying to address.

objective correctness
self-judgement
hidden-state readouts
semantic ambiguity
transferability
Innovation

Methods, ideas, or system contributions that make the work stand out.

self-judgement confounding
objective correctness
hidden-state readouts
transferability asymmetry
directional probing