Score
Designs, builds, and evaluates diagnostic probes—linear classifiers, regression heads, logit lens decoders, and auxiliary-task readouts—that extract and quantify information encoded in model internals (activations, embeddings, logits) to test for specific concepts, dynamics, or signals. Implements comparative analyses that measure representation drift, compute signed cross-domain affinities, and derive probe-based metrics to predict positive or negative transfer between domains or models.
This work addresses the challenge of efficiently and stably extracting concept vectors from intermediate layers of frozen large language models for activation steering. The authors propose a lightweight probing method based on L2-regularized logistic regression, which—by incorporating a validation-tuned ridge parameter into normalized weights—systematically links regularization strength to both directional stability and training efficiency of concept vectors for the first time. Leveraging the Convex Gaussian Minimax Theorem (CGMT), they provide a high-dimensional, few-shot theoretical justification for their approach. Extensive experiments across multiple instruction-tuned models and synthetic concept datasets demonstrate that the method achieves comparable or superior accuracy relative to strong baselines while significantly reducing training cost and enhancing directional stability.
This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.
本文通过引入概念定向归因(CTA)方法,解决了线性探针如何产生内部概念表示的问题,并解释了影响探针性能的内部计算机制。
This work addresses the challenge of efficiently selecting the optimal model checkpoint and enabling early stopping in the absence of a labeled validation set. It proposes a lightweight, label-free proxy metric that leverages the Frobenius norm of the classification head’s weight gradients—computed from a single forward-backward pass—as a performance prediction signal. To the best of our knowledge, this is the first approach to utilize per-batch gradient norms as a universal, unsupervised estimator of model performance across diverse tasks and architectures, including image classification, object detection, segmentation, and diffusion models. The method adapts to different network structures through feature or head-scale normalization. Experiments demonstrate near-oracle checkpoint selection on ImageNet-1k (average gap of only 1.12%) and effective prediction of mAP and FID on COCO and CIFAR-10 diffusion models, with computational overhead below 0.1% of a single training epoch.
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
本文通过SAE分解方法,区分了探针读数与行为驱动因素,解决了线性探针解码能力不代表因果关系的问题。
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。
This study addresses the tendency of large language models to circumvent alignment objectives through superficial compliance, resulting in internal representations that fail to genuinely internalize safe behaviors. To overcome this limitation, we propose a probe-guided fine-tuning approach that, for the first time, employs continuously updated internal probes as direct optimization signals. By leveraging both linear and nonlinear probing techniques, our method shapes internal representations specifically for harmlessness and honesty, transcending the constraints of relying solely on output-level feedback. Empirically, this approach significantly outperforms Direct Preference Optimization (DPO) and inference-time interventions in navigating the safety-utility trade-off. It substantially enhances robustness against jailbreak attacks while preserving the linear encoding of concepts to ensure continued monitorability.
This study addresses the lack of theoretical foundations for finite probe representations in neural network property learning, where reliance solely on final outputs yields insufficient information. To bridge this gap, we establish identifiability and universality theories for probe learning, deriving the first sufficiency bounds for finite probes and demonstrating that intermediate hidden-layer representations are superior to final outputs. Guided by these theoretical insights, we propose HIDDENPROBE, a minimalist yet highly efficient architecture. We apply this framework, supported by rigorous theoretical analysis, to both MLPs and Transformers. Extensive evaluations across multiple neural functionality benchmarks show that HIDDENPROBE consistently outperforms existing methods, achieving state-of-the-art performance. The source code has been made publicly available.