layer-wise probing

Design and run per-layer diagnostic probes and causal interventions on a model's hidden activations to measure, localize, and compare where and how strongly specific information is represented across network depth. Typical work includes training layer-specific (often linear) classifiers, performing activation perturbations or intervention probes and layer sweeps, and reporting per-layer accuracy, robustness, emergence patterns, and decomposed signal contributions.

layer-wiseprobing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.02
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$212K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.

Comparing zero-shot baselines with probing methods for efficiencyImproving activation monitoring via prompted probing and sparse autoencodersMonitoring language model outputs for unexpected unsafe behaviors

Existing neural network probing methods often rely on input perturbations or parameter analysis, which struggle to uncover structured information embedded in intermediate representations. This work proposes APEX, a novel probing paradigm that perturbs hidden activations during inference while keeping both inputs and model parameters fixed. APEX formalizes activation perturbation as a general probing framework, unifying and extending prior approaches such as input perturbation as special cases, and enabling a controllable transition from sample-dependent to model-dependent behavioral analysis. Experiments demonstrate that APEX effectively quantifies representational structure, distinguishes models trained on structured versus random labels, reveals semantically coherent prediction transitions, and precisely identifies the concentration of predictions toward target classes in backdoor attacks.

activation perturbationintermediate representationsneural networks

This work addresses the fragility of single-layer linear probes in detecting "deliberate" erroneous outputs from language models, where the optimal layer varies across models and tasks. To enhance robustness, the authors propose an ensemble of multi-layer linear probes that leverages the geometric property of deception directions progressively rotating across model layers. Systematic evaluation across models ranging from 0.5B to 176B parameters demonstrates that the multi-layer ensemble improves AUROC by 29% on the Insider Trading task and by 78% on the Harm-Pressure Knowledge task. Furthermore, the study quantifies—for the first time—the scaling behavior of probe performance with model size, revealing that AUROC increases by approximately 5% per tenfold increase in parameter count (R = 0.81).

deception detectionlinear probemodel scaling

Deep Linear Probe Generators for Weight Space Learning

Oct 14, 2024
JK
Jonathan Kahana
🏛️ The Hebrew University of Jerusalem

Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.

Addressing ineffectiveness of current weight space probing methodsImproving probe learning strategies for neural network analysisReducing computational costs while maintaining model performance

How Reliable are Causal Probing Interventions?

Aug 28, 2024
ME
Marc E. Canby
🏛️ University of Illinois Urbana-Champaign

This work addresses the lack of rigorous reliability evaluation for causal probing interventions in large language models. We propose the first quantifiable and comparable two-dimensional empirical framework, formalizing intervention effectiveness via “completeness” and “selectivity,” and defining their harmonic mean as the core “reliability” metric. Through hierarchical, controlled intervention experiments and cross-method benchmarking, we formally uncover fundamental reliability trade-offs: no single method achieves universal reliability across all network layers; nonlinear interventions outperform linear ones in shallow-to-middle layers, whereas linear interventions exhibit greater robustness in deeper layers; and concept removal methods are significantly less reliable than counterfactual interventions—challenging their validity for causal explanation. Our framework establishes a theoretical benchmark and practical guidelines for causal interpretability research in foundation models.

Analyze tradeoff between completeness and selectivityCompare effectiveness of different intervention familiesEvaluate reliability of causal probing methods

Latest Papers

What's happening recently
View more

This study addresses the lack of theoretical foundations for finite probe representations in neural network property learning, where reliance solely on final outputs yields insufficient information. To bridge this gap, we establish identifiability and universality theories for probe learning, deriving the first sufficiency bounds for finite probes and demonstrating that intermediate hidden-layer representations are superior to final outputs. Guided by these theoretical insights, we propose HIDDENPROBE, a minimalist yet highly efficient architecture. We apply this framework, supported by rigorous theoretical analysis, to both MLPs and Transformers. Extensive evaluations across multiple neural functionality benchmarks show that HIDDENPROBE consistently outperforms existing methods, achieving state-of-the-art performance. The source code has been made publicly available.

identifiabilityneural functionalsprobe-based representations

This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.

in-context decodinginformation-theoretic decodabilitymodel errors

Hot Scholars

FS

Fei Shen

National University of Singapore
Controllable GenerationMultimodal Safety
YB

Yonatan Belinkov

Technion
Natural Language ProcessingModel InterpretabilityArtificial Intelligence
EN

Ercong Nie

LMU Munich, MCML
Computational LinguisticsNatural Language Processing
VS

Vasu Sharma

Facebook AI Research (FAIR)
Generative AILLMsComputer VisionNatural Language Processing
LH

Lijie Hu

Assistant Professor, MBZUAI
Explainable AILLMDifferential Privacy