π€ AI Summary
Existing correctness probes struggle to disentangle objective correctness (OC) from the modelβs self-judgment (SJ) in language model hidden states due to their high empirical alignment. This work constructs samples where OC and SJ diverge and employs a factorization-based approach to isolate and estimate the directions associated with each signal, evaluating their polarity and transferability across mathematical reasoning and factual recall tasks. Experiments on four instruction-tuned models (up to 14B parameters) reveal that current probes predominantly capture SJ rather than OC; furthermore, the SJ-associated direction exhibits strong cross-task transferability, whereas the OC direction performs worse than random. These findings systematically demonstrate, for the first time, a fundamental limitation in the semantic reliability of prevailing correctness probes.
π Abstract
Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.