Score
Design and run per-layer diagnostic probes and causal interventions on a model's hidden activations to measure, localize, and compare where and how strongly specific information is represented across network depth. Typical work includes training layer-specific (often linear) classifiers, performing activation perturbations or intervention probes and layer sweeps, and reporting per-layer accuracy, robustness, emergence patterns, and decomposed signal contributions.
This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.
Existing neural network probing methods often rely on input perturbations or parameter analysis, which struggle to uncover structured information embedded in intermediate representations. This work proposes APEX, a novel probing paradigm that perturbs hidden activations during inference while keeping both inputs and model parameters fixed. APEX formalizes activation perturbation as a general probing framework, unifying and extending prior approaches such as input perturbation as special cases, and enabling a controllable transition from sample-dependent to model-dependent behavioral analysis. Experiments demonstrate that APEX effectively quantifies representational structure, distinguishes models trained on structured versus random labels, reveals semantically coherent prediction transitions, and precisely identifies the concentration of predictions toward target classes in backdoor attacks.
This work addresses the fragility of single-layer linear probes in detecting "deliberate" erroneous outputs from language models, where the optimal layer varies across models and tasks. To enhance robustness, the authors propose an ensemble of multi-layer linear probes that leverages the geometric property of deception directions progressively rotating across model layers. Systematic evaluation across models ranging from 0.5B to 176B parameters demonstrates that the multi-layer ensemble improves AUROC by 29% on the Insider Trading task and by 78% on the Harm-Pressure Knowledge task. Furthermore, the study quantifies—for the first time—the scaling behavior of probe performance with model size, revealing that AUROC increases by approximately 5% per tenfold increase in parameter count (R = 0.81).
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
This work addresses the lack of rigorous reliability evaluation for causal probing interventions in large language models. We propose the first quantifiable and comparable two-dimensional empirical framework, formalizing intervention effectiveness via “completeness” and “selectivity,” and defining their harmonic mean as the core “reliability” metric. Through hierarchical, controlled intervention experiments and cross-method benchmarking, we formally uncover fundamental reliability trade-offs: no single method achieves universal reliability across all network layers; nonlinear interventions outperform linear ones in shallow-to-middle layers, whereas linear interventions exhibit greater robustness in deeper layers; and concept removal methods are significantly less reliable than counterfactual interventions—challenging their validity for causal explanation. Our framework establishes a theoretical benchmark and practical guidelines for causal interpretability research in foundation models.
本文通过引入概念定向归因(CTA)方法,解决了线性探针如何产生内部概念表示的问题,并解释了影响探针性能的内部计算机制。
This study addresses the lack of theoretical foundations for finite probe representations in neural network property learning, where reliance solely on final outputs yields insufficient information. To bridge this gap, we establish identifiability and universality theories for probe learning, deriving the first sufficiency bounds for finite probes and demonstrating that intermediate hidden-layer representations are superior to final outputs. Guided by these theoretical insights, we propose HIDDENPROBE, a minimalist yet highly efficient architecture. We apply this framework, supported by rigorous theoretical analysis, to both MLPs and Transformers. Extensive evaluations across multiple neural functionality benchmarks show that HIDDENPROBE consistently outperforms existing methods, achieving state-of-the-art performance. The source code has been made publicly available.
本文通过SAE分解方法,区分了探针读数与行为驱动因素,解决了线性探针解码能力不代表因果关系的问题。
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
研究使用线性探针检测大型语言模型中工具调用错误的有效性,通过18个模型测试,发现该方法能有效识别多种错误。