Score
Designs, trains, and evaluates linear classifiers (linear probes) on frozen or pooled feature representations to diagnose and quantify which aspects of the representation are linearly decodable, to discover informative linear directions, and to compare linear versus nonlinear decodability. This includes constructing and analyzing low‑rank, rank‑constrained, compressible, pooled, self‑supervised, and pretrained/transferable probes and measuring probe capacity, sample efficiency, generalization across distributions, and performance as a function of rank to inform lightweight decoder design.
Direct inference of training/generalization error from model weights remains challenging due to high dimensionality and neuron permutation symmetry in weight-space learning. Method: We propose ProbeGen—a deep linear probe generator that introduces a shared, deep linear generative module to inject structural inductive bias into input probes, thereby substantially mitigating overfitting inherent in conventional probe learning. By analyzing the output responses of structured probes via forward propagation, ProbeGen achieves efficient representation of the weight space. Contribution/Results: Across multiple benchmarks, ProbeGen outperforms state-of-the-art methods with 30–1000× lower computational cost (significantly reduced FLOPs) and enhanced robustness. To our knowledge, this is the first work to systematically integrate structured probe generation with weight-space representation learning, establishing a novel paradigm for model diagnosis and generalization analysis.
Assessing the statistical reliability of linear classifiers in high-dimensional biomedical data—such as ER-positive breast cancer gene expression profiles—remains challenging, particularly in distinguishing true discriminative power from spurious performance due to random correlations. Method: We propose a novel homogeneity test grounded in linear separability, introducing the first analytically derived upper bound on the linear separability *p*-value. We rigorously prove its high accuracy under bivariate normality and extend it to high dimensions, enabling strict statistical inference on classifier significance—not merely random efficacy. The method integrates linear separability analysis, derivation of *p*-value upper bounds, and paired gene expression testing. Contribution/Results: Applied to ER-positive breast cancer recurrence prediction, our approach identifies the IGFBP6–ELOVL5 gene pair as exhibiting statistically significant synergistic discriminative capacity (*p* < 0.01), offering a new paradigm for interpretable biomarker discovery.
In scientific machine learning, modeling the mapping from physical processes to observed data faces dual challenges of interpretability and rank deficiency. This paper establishes a unified theoretical framework for linear encoder-decoder architectures grounded in Bayesian risk minimization, and— for the first time—derives closed-form optimal linear/affine mappings under explicit rank constraints, systematically addressing low-rank degeneracies in data, operators, and measurements. Our approach integrates Bayesian decision theory with low-rank matrix optimization, avoiding black-box nonlinear models. Evaluated on biomedical imaging, financial factor analysis, and shallow-water equation simulation, the proposed linear baseline achieves strong interpretability, high reproducibility, and superior generalization. It thus provides a trustworthy, robust, and benchmarkable paradigm for scientific AI.
This work addresses the limitation of conventional cosine similarity in evaluating linear probes, which neglects task-specific data distributions and thus fails to reliably predict out-of-distribution (OOD) performance. The authors propose Mahalanobis Cosine Similarity (MCS), a task-aware probe similarity metric that incorporates the covariance of test data to weight the inner product. Under assumptions of Gaussian projections and class balance, they theoretically establish—for the first time—that both OOD AUROC and MCS are sigmoidal functions of the signal-to-noise ratio, yielding an approximate linear relationship between them and delineating the boundary conditions under which this relationship breaks down. Extensive experiments across diverse models, network layers, and concept domains demonstrate a strong linear correlation (R² = 0.98) between MCS and OOD AUROC, substantially outperforming standard cosine similarity.
Existing single-view linear probes struggle to model the higher-order interaction structures between rows and columns in model weights, limiting the effectiveness of weight space learning. To address this, this work proposes MVProbe—the first multi-view probing framework tailored for weight representations—which explicitly captures higher-order correlations by fusing first-order signals with an interaction-aware view constructed via Gram matrices. The method introduces learnable probe vectors and incorporates a scaling-law-guided normalization strategy to enable adaptive normalization and fusion of multi-branch features. Evaluated on the Model Jungle benchmark, MVProbe consistently outperforms the current state-of-the-art ProbeX across diverse architectures, including ResNet, SupViT, MAE, DINO, and Stable Diffusion LoRA.
This study investigates whether representations captured by linear probes in the hidden states of large language models stem from genuine differences in reasoning mechanisms or are confounded by task formatting. Focusing on Qwen3-14B across three reasoning tasks, we conduct a systematic analysis integrating linear probing, residualization-based controls, intrinsic dimension estimation, convex hull contamination analysis, trajectory anchor similarity, and causal intervention experiments. Our results show that while raw probe accuracy reaches 100%, it drops to chance level after controlling for task format. Causal testing further reveals no significant functional association between the geometric structure of hidden states and reasoning patterns (p = 0.286). These findings demonstrate that standard probing analyses are highly susceptible to format confounding and underscore the necessity of routinely incorporating format-deconfounding controls in mechanistic interpretability research.
This study addresses the lack of theoretical foundations for finite probe representations in neural network property learning, where reliance solely on final outputs yields insufficient information. To bridge this gap, we establish identifiability and universality theories for probe learning, deriving the first sufficiency bounds for finite probes and demonstrating that intermediate hidden-layer representations are superior to final outputs. Guided by these theoretical insights, we propose HIDDENPROBE, a minimalist yet highly efficient architecture. We apply this framework, supported by rigorous theoretical analysis, to both MLPs and Transformers. Extensive evaluations across multiple neural functionality benchmarks show that HIDDENPROBE consistently outperforms existing methods, achieving state-of-the-art performance. The source code has been made publicly available.
This work addresses the limitation of existing neural classifiers that rely on linear readouts and struggle to capture the geometric structure of class representations, particularly under few-shot conditions where unilateral affine separability cannot be properly assessed. The authors propose a directional Linear Separability Metric (LSM) that quantifies the minimal proportion of competing-class samples intruding into an affine half-space containing all samples of a target class. LSM exhibits asymmetry, class-level granularity, target normalization, and invariance under full-rank linear transformations, thereby distinguishing the effects of linear reparameterizations from those of information loss or nonlinear distortions. An efficient penalty-based affine search algorithm is introduced to estimate LSM in high-dimensional feature spaces while preserving the original discrete constraints. Experiments demonstrate that LSM effectively reveals class intrusion phenomena induced by components such as coordinate gating, offering a novel tool for analyzing the geometry of neural representations.
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
This work addresses the round complexity of boosting algorithms for concept classes satisfying XOR closure. Classical boosting requires $O(\log(1/\varepsilon)/\gamma^2)$ calls to a weak learner to achieve error $\varepsilon$, matching known theoretical lower bounds. The authors establish, for the first time, a connection between boosting and list-decodable codes for such concept classes: by treating the target function as an encodable message and leveraging the weak learner’s outputs together with list decoding techniques, they efficiently identify a strong hypothesis from a candidate set. This approach breaks the classical round-complexity barrier, achieving the desired accuracy with only $O(\log(1/\varepsilon))$ rounds of weak learner queries, supplemented by $\tilde{O}(\log(1/\varepsilon)/\gamma^2)$ additional samples for verification, thereby significantly improving learning efficiency.