When Do Internal Probes Beat Reading the Answer? Miscalibrated Readouts and Behavior-Concealed Knowledge in Language Models

📅 2026-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了语言模型在逻辑验证任务中行为与内部状态不一致的问题,通过线性探针分析隐藏状态,并提出了一种单参数校正方法来修复模型的行为。
📝 Abstract
A 0.6B language model, asked to verify 1,200 logical conclusions (half valid, half corrupted by a single semantic edit), answers YES every time. Judged by behavior it discriminates nothing; linear probes on its hidden states read the correct verdict at 0.96 AUC, transferring to unseen logical structures and separating foils built from exactly the words of the true conclusion (0.90). We ask where the verdict is lost, and find the dominant failure is a single scalar. The verdict survives to the model's own output logits (margin AUC 0.89) along a well-aligned readout direction; a saturated decision threshold, offset by +4.6 sigma, erases it. The diagnosis generalizes: across 90 semantic-label configurations of a five-model, three-family factorial, behavioral accuracy collapses onto a single function of threshold offset (Spearman -0.93) while margin ranking moves far less. Across a 13x scale range, internal knowledge saturates while free-form behavior is non-monotone: an 8B model underperforms its 4B sibling through an answer-channel failure rather than the threshold; forced-choice accuracy is monotone. The diagnosis is actionable: a one-parameter correction, never fit on evaluated structures, repairs behavior from 50% to 81% (0.6B); calibrated margin decoding recovers 94% at 8B; few-shot prompting works the same way, recentering the threshold (+4.6 sigma to 0.0 sigma) while preserving ranking. Comparing probe to margin separates three regimes: concealed, miscalibrated, and undetected. On a maze task built so foils carry no surface cues, the audit correctly reports the third. In the standard generation setting, answer-surface features and heuristic labels reproduce published probing results without any internal access.
Problem

Research questions and friction points this paper is trying to address.

language model
logical conclusions
hidden states
behavioral accuracy
threshold offset
Innovation

Methods, ideas, or system contributions that make the work stand out.

internal probes
miscalibrated readouts
hidden states
decision threshold
behavior-concealed knowledge
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gnaneswar Villuri
Department of Electrical and Computer Engineering, Stony Brook University, Stony Brook, NY 11794, USA
Hashmath Shaik
Hashmath Shaik
Research Assistant
AIMachine LearningDeep Learning
A
Alex Doboli
Department of Electrical and Computer Engineering, Stony Brook University, Stony Brook, NY 11794, USA