Score
Design, train, and evaluate lightweight probe models that read internal model states (e.g., hidden activations, logits, attention patterns) to estimate prediction uncertainty or the likelihood of error/hallucination. This work includes choosing probe architectures and objectives, constructing labels and prompts for supervision, and comparing probe designs and calibration under matched conditions.
Existing probing-based uncertainty estimation methods suffer from entangled design choices in feature representation, training data, and evaluation protocols, making it difficult to isolate the key drivers of performance improvements. This work presents the first factorized dissection of probing approaches, systematically evaluating—within a unified framework—the impact of signal types, prompting strategies, and label construction schemes. We introduce a transferable pre-trained probe that leverages hidden states and attention features, followed by structured compression and cross-task transfer training. Our analysis reveals that raw hidden states excel in in-domain settings, whereas structured features demonstrate superior robustness under distributional shift. The proposed method achieves effective transfer on open-ended fact generation tasks, establishing a stable and reliable baseline for uncertainty estimation.
This study investigates how dataset suitability and large language model (LLM) response uncertainty affect probe model performance. We propose a “response uncertainty–feature interpretability” analytical framework, empirically establishing for the first time a strong negative correlation between LLM output entropy/variance and probe accuracy. To attribute uncertainty sources, we introduce a gradient- and attention-based uncertainty attribution mechanism that quantifies feature importance. Furthermore, we evaluate LLM internal representations against human knowledge using a multi-task interpretability benchmark. Results show that reducing response uncertainty significantly improves probe performance; moreover, high-consistency reasoning instances—identified via our framework—exhibit robust cross-task and cross-domain stability. These findings offer a novel pathway toward trustworthy and interpretable AI.
This study investigates whether representations captured by linear probes in the hidden states of large language models stem from genuine differences in reasoning mechanisms or are confounded by task formatting. Focusing on Qwen3-14B across three reasoning tasks, we conduct a systematic analysis integrating linear probing, residualization-based controls, intrinsic dimension estimation, convex hull contamination analysis, trajectory anchor similarity, and causal intervention experiments. Our results show that while raw probe accuracy reaches 100%, it drops to chance level after controlling for task format. Causal testing further reveals no significant functional association between the geometric structure of hidden states and reasoning patterns (p = 0.286). These findings demonstrate that standard probing analyses are highly susceptible to format confounding and underscore the necessity of routinely incorporating format-deconfounding controls in mechanistic interpretability research.
This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
This study systematically evaluates the generalization and reproducibility of lightweight safety probes across diverse large language model families. Building upon final-layer MLP activations, we reproduce and extend the original probing methodology on prominent models—including LLaMA, Gemma, Mistral, and Qwen2—and assess its effectiveness in detecting harmful inputs using three established safety benchmarks: WildJailbreak, BeaverTails, and AEGIS 2.0. Our experiments demonstrate that the probe achieves consistent performance across all evaluated models, with F1 score variations under 1 percentage point and a reproduction error below 0.37%. Notably, we uncover for the first time that the representation of the final token exhibits invariance to inference-time random seeds, thereby confirming the robustness and cross-model applicability of this probing approach.
This work addresses the low accuracy—under 50%—of large language models (LLMs) in filling tool parameters within complex domains such as cloud networking. The authors propose a unified framework that, for the first time, reveals strong correctness signals embedded in LLM hidden states. Leveraging this insight, they construct a linear probe and integrate it into a novel pipeline featuring probe-guided bootstrapped training (PBT) and probe-guided re-ranking (PGR) during inference. Additionally, they introduce ParamBench, the first benchmark annotated by parameter nesting depth, inter-parameter dependencies, and reasoning complexity. Evaluated on ParamBench and six external benchmarks, the proposed method substantially improves the average exact match accuracy of five open-source LLMs from 19.7% to 59.6%.