Score
Design and train lightweight diagnostic probes—typically small classifiers or regressors applied to frozen decoder activations at specific layers and time steps—to read out information from a model’s intermediate decoder states. Build layerwise and timestep-wise analyses and evaluation protocols that identify which internal representations encode particular signals (e.g., next-token probabilities, error or hallucination indicators) and quantify how those signals evolve through the decoder.
This paper addresses the inefficiency and delayed detection of unsafe behaviors in large language models (LLMs) due to coarse-grained internal activation monitoring. We propose a task-prompt-driven collaborative monitoring framework. Methodologically, we conduct the first systematic comparison of three techniques: zero-shot prompting, prompt-augmented linear probing, and sparse autoencoder (SAE)-based activation pooling; integrate task-specific prompt engineering with SAE representation learning; and employ token-level max-pooling to enhance signal robustness. Key contributions include: (1) demonstrating that prompt-based probing achieves superior data efficiency and cross-task generalization compared to alternatives; (2) showing that SAE-enhanced probing outperforms raw activation monitoring under low inference overhead; and (3) establishing optimal monitoring paradigms under varying computational constraints—providing both theoretical foundations and practical guidelines for efficient, deployable LLM safety monitoring.
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
Traditional decoding probes struggle to independently assess the contribution of individual features to language model representations and are susceptible to confounding effects arising from feature correlations. This work proposes a novel encoding probe paradigm that analyzes the information encoded in model representations by linearly reconstructing internal states from interpretable features—such as acoustic, phonological, syntactic, lexical, and speaker identity attributes. This approach enables, for the first time, disentangled evaluation of each feature’s contribution and facilitates cross-modal comparisons. Experimental results demonstrate that syntactic and lexical features provide independent contributions to representation reconstruction, whereas speaker-related effects are highly dependent on training objectives and datasets, thereby validating the method’s efficacy and complementarity to existing probing techniques.
This work addresses the challenge of efficiently selecting the optimal model checkpoint and enabling early stopping in the absence of a labeled validation set. It proposes a lightweight, label-free proxy metric that leverages the Frobenius norm of the classification head’s weight gradients—computed from a single forward-backward pass—as a performance prediction signal. To the best of our knowledge, this is the first approach to utilize per-batch gradient norms as a universal, unsupervised estimator of model performance across diverse tasks and architectures, including image classification, object detection, segmentation, and diffusion models. The method adapts to different network structures through feature or head-scale normalization. Experiments demonstrate near-oracle checkpoint selection on ImageNet-1k (average gap of only 1.12%) and effective prediction of mAP and FID on COCO and CIFAR-10 diffusion models, with computational overhead below 0.1% of a single training epoch.
This work addresses the challenge of efficiently identifying optimal configurations in 3D-CT vision-language models, where combining frozen image encoders with token compression schemes typically requires costly fine-tuning. To circumvent this, the authors propose a low-cost probing method that predicts downstream fine-tuning performance directly from cached embeddings of frozen encoders. They introduce an image-guided probing benchmark to evaluate the ranking efficacy of various (encoder × compression) combinations and incorporate two novel validation mechanisms—scale-sanity and probe-separability—to ensure clinically relevant attributes are decodable and representation scales remain reasonable. Experimental results demonstrate that the probe-based rankings correlate strongly with full fine-tuning outcomes (Spearman’s ρ ≈ 0.95), enabling candidate configuration screening within minutes and substantially reducing computational overhead.
This work addresses the lack of scalable security auditing mechanisms in existing AI coding agents, particularly the difficulty of vulnerability detection in closed-source models. The authors propose training linear probes on intermediate activation layers of open-source large language models to directly extract security signals from internal representations, enabling discrimination between vulnerable and patched code functions. They demonstrate for the first time that internal model activations contain security-relevant information not captured by output-based prompting, and that a single linear probe generalizes effectively to unseen vulnerability types. Evaluated across five mainstream models, the approach achieves 61–67% accuracy—significantly outperforming output-prompting baselines (~50%)—and remains robust regardless of prompting strategies such as chain-of-thought.
This work addresses the inefficiency of conventional content moderation for large language models deployed on user devices, where post-generation filtering with a separate model doubles inference costs and precludes real-time intervention. The authors propose leveraging intrinsic safety signals embedded in the model’s hidden states to train lightweight, token-level linear probes that assess output safety during decoding—without requiring additional forward passes. This approach enables streaming content moderation through early termination or correction during generation, supported by dynamic thresholds and token-level score aggregation. Remarkably, a single probe at an intermediate layer replicates most decisions of a strong standalone moderator while reducing computational overhead by several orders of magnitude, achieving sub-millisecond per-token safety checks with minimal latency and cost.