Score
Design, build, and evaluate mechanisms that read out latent variables or task-relevant signals from neural network intermediate activations (e.g., residual stream or layer hidden states) by extracting hidden-state features and mapping them to targets via lightweight probes, linear regressors/classifiers, clustering, or matching methods; include training-free readouts and multi-target inference, layerwise linear separability diagnostics, and assessments of probe generalization across models and tasks.
This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.
This study investigates whether representations captured by linear probes in the hidden states of large language models stem from genuine differences in reasoning mechanisms or are confounded by task formatting. Focusing on Qwen3-14B across three reasoning tasks, we conduct a systematic analysis integrating linear probing, residualization-based controls, intrinsic dimension estimation, convex hull contamination analysis, trajectory anchor similarity, and causal intervention experiments. Our results show that while raw probe accuracy reaches 100%, it drops to chance level after controlling for task format. Causal testing further reveals no significant functional association between the geometric structure of hidden states and reasoning patterns (p = 0.286). These findings demonstrate that standard probing analyses are highly susceptible to format confounding and underscore the necessity of routinely incorporating format-deconfounding controls in mechanistic interpretability research.
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
This work addresses the challenge of efficiently selecting the optimal model checkpoint and enabling early stopping in the absence of a labeled validation set. It proposes a lightweight, label-free proxy metric that leverages the Frobenius norm of the classification head’s weight gradients—computed from a single forward-backward pass—as a performance prediction signal. To the best of our knowledge, this is the first approach to utilize per-batch gradient norms as a universal, unsupervised estimator of model performance across diverse tasks and architectures, including image classification, object detection, segmentation, and diffusion models. The method adapts to different network structures through feature or head-scale normalization. Experiments demonstrate near-oracle checkpoint selection on ImageNet-1k (average gap of only 1.12%) and effective prediction of mAP and FID on COCO and CIFAR-10 diffusion models, with computational overhead below 0.1% of a single training epoch.
Understanding the causal roles and computational mechanisms of intermediate variables—such as syntactic attributes—in Transformer language models remains challenging. Method: We propose *circuit probing*, a hypothesis-driven methodology that reverse-engineers intermediate representations encoding specific linguistic properties and precisely identifies the parameter-level neural circuits supporting them. Our approach integrates gradient-guided discovery, targeted parameter ablation, verification via diagnostic probe training, and modular attribution analysis to enable causal intervention and algorithm-level interpretation. Contribution/Results: This work unifies hypothesis testing, circuit localization, and dynamic tracing for the first time, revealing implicit algorithmic structures within models and their training-time evolution. Experiments successfully decode symbolic arithmetic logic in specialized arithmetic models, localize subject–verb agreement and reflexive pronoun processing circuits in GPT-2, and empirically confirm their progressive emergence during training.
This work addresses the inefficiency of conventional content moderation for large language models deployed on user devices, where post-generation filtering with a separate model doubles inference costs and precludes real-time intervention. The authors propose leveraging intrinsic safety signals embedded in the model’s hidden states to train lightweight, token-level linear probes that assess output safety during decoding—without requiring additional forward passes. This approach enables streaming content moderation through early termination or correction during generation, supported by dynamic thresholds and token-level score aggregation. Remarkably, a single probe at an intermediate layer replicates most decisions of a strong standalone moderator while reducing computational overhead by several orders of magnitude, achieving sub-millisecond per-token safety checks with minimal latency and cost.
It remains unclear whether existing training-based interpretability methods for neural networks can reveal information beyond what is present in their training data. To address this, this work proposes HARP, the first fully training-free approach that achieves high-performance interpretability by equipping large language model agents with an activation vector database and tools such as directional projection and activation differencing, enabling hypothesis-driven, on-demand retrieval and probing. HARP outperforms activation oracles and SAE-based agents across tasks including concept discovery, detection, model steering, and secret extraction, while flexibly indexing new datasets. These results demonstrate that current training-based methods have not yet surpassed the informational limits of their training data and establish HARP as a more efficient and cost-effective alternative.
This work addresses the low accuracy—under 50%—of large language models (LLMs) in filling tool parameters within complex domains such as cloud networking. The authors propose a unified framework that, for the first time, reveals strong correctness signals embedded in LLM hidden states. Leveraging this insight, they construct a linear probe and integrate it into a novel pipeline featuring probe-guided bootstrapped training (PBT) and probe-guided re-ranking (PGR) during inference. Additionally, they introduce ParamBench, the first benchmark annotated by parameter nesting depth, inter-parameter dependencies, and reasoning complexity. Evaluated on ParamBench and six external benchmarks, the proposed method substantially improves the average exact match accuracy of five open-source LLMs from 19.7% to 59.6%.