Score
Designs and implements methods to estimate which low-dimensional linear subspaces of neural network activations are consistently identifiable across random seeds, mini‑batches, prompts, or training and evaluation conditions. This work builds pipelines and metrics to compute top‑r subspaces, quantify reproducible directions and fraction of shared basis, and compare subspaces using alignment measures such as chordal distance, basis overlap, CKA/RSA/SVCCA, activation‑delta comparisons, and spectral‑tail scaling analyses.
This paper systematically investigates the quantification of neural network model similarity, addressing both representational similarity (intermediate-layer activations) and functional similarity (output behavior). It unifies and comparatively analyzes mainstream metrics—such as CKA, SVCCA, PWCCA, linear probes, and top-k output agreement—across these two complementary paradigms. A structured taxonomy is introduced, accompanied by theoretical analysis of each metric’s mathematical properties, interrelationships, and applicability boundaries. Empirical evaluation assesses their explanatory power and limitations in downstream tasks including model compression, ensemble learning, and robustness analysis. The core contribution is a cross-paradigm benchmark enabling rigorous metric comparison, revealing how metric selection critically influences downstream conclusions. The work further identifies open challenges and proposes principled evaluation criteria, thereby advancing methodological foundations for model behavior interpretation and trustworthy AI. (149 words)
This work investigates whether deep neural networks trained on diverse tasks, initializations, and domains share a common low-dimensional parameter subspace. Method: We perform systematic low-dimensional subspace extraction and comparison across over 1,100 model weight matrices—spanning heterogeneous architectures, tasks, and domains—using per-modal spectral analysis and singular value decomposition (SVD). Contribution/Results: We provide the first empirical evidence that, despite substantial differences in training configurations, models consistently converge to highly overlapping low-dimensional weight subspaces. Remarkably, fewer than 1% of the original parameter dimensions—captured by the top principal directions—are sufficient to explain over 90% of the parameter variance. This universal subspace exhibits strong reusability, enabling data-efficient parameter structure discovery, scalable multi-task learning, model fusion, and compression—thereby significantly reducing training cost and computational resource requirements.
Representational similarity metrics (e.g., CCA, CKA) systematically underestimate model–neuron alignment when neuron counts are limited, hindering reliable inference in computational neuroscience. Method: We propose a spectral analysis framework grounded in random matrix theory, establishing the first theoretical characterization of how finite sampling biases spectral estimates of CCA/CKA—specifically revealing that eigenvector delocalization induces systematic underestimation. Building on this insight, we design a spectral denoising correction method enabling unbiased population-level similarity estimation from small samples (<100 neurons). Contribution/Results: Validated on synthetic data and multiple real neural datasets (e.g., V4, IT cortex), our approach significantly improves estimation accuracy and interpretability of representational similarity. It provides both theoretical foundations and a practical tool for model evaluation under sparse neural recording conditions, advancing small-sample–driven computational neuroscience.
This work investigates whether semantic factors in neural network representation spaces can be unsupervisedly decomposed into interpretable, orthogonal subspaces. We propose Neighborhood Distance Minimization (NDM), the first fully unsupervised method—without basis alignment assumptions—to learn class-variable-directed, interpretable subspaces, revealing structured, “functional-circuit”-like organization within model internals. Qualitative analysis and quantitative validation against known GPT-2 circuits confirm strong correlations between discovered subspaces and specific semantic variables (e.g., grammatical roles, factual knowledge). Furthermore, we successfully separate contextual representation from knowledge routing in a 2-billion-parameter GPT-2 model, demonstrating the method’s scalability and practical utility. Our approach advances interpretability by enabling decomposition of high-dimensional representations into semantically meaningful, disentangled subspaces without supervision or architectural constraints.
Existing activation alignment methods struggle to capture differences in the sensitivity of neural representations to local stimulus perturbations and thus fail to reflect how systems leverage local evidence for discrimination. This work proposes a novel analytical framework based on locally decodable information, integrating Fisher information, pullback metrics, and log-spectral distances on the SPD manifold to construct the Spectral Riemannian Alignment Score (S-RAS). S-RAS provides, for the first time, a minimal, dataset-level summary of neural representational sensitivity from the perspective of local discriminative tasks, with guaranteed multiplicative consistency. The method successfully aligns corresponding layers across independently trained networks, enables transferable class-conditional probing, reveals representational differences between standard and robustly trained models, and uncovers stimulus coordinate family effects in mouse visual cortex.
This study addresses the challenge of comparing high-dimensional neural representations across neuroscience and artificial intelligence: specifically, how to select similarity measures that best reveal functional correspondences and divergences. We systematically evaluate eight mainstream representational similarity metrics—including linear CKA, Procrustes distance, CCA, inner-product kernel, and nearest-neighbor alignment—against behavioral functional alignment (e.g., recognition accuracy, generalization, robustness) as a ground-truth benchmark. Our evaluation spans both biological neural data and artificial neural network models. Results show that geometry-sensitive metrics—particularly linear CKA and Procrustes distance—consistently outperform predictive metrics, achieving superior alignment with human behavioral performance and effectively distinguishing trained versus untrained models. In contrast, linear predictivity exhibits only moderate behavioral alignment. This work establishes the first behavior-driven representational similarity benchmark, providing a principled, cross-domain methodology for mechanistic interpretation and comparative analysis of neural computation.
Existing representational alignment methods—such as Representational Similarity Analysis (RSA), Centered Kernel Alignment (CKA), and linear regression—systematically underestimate true similarity when neural representations reside in superposition. This bias stems from conflating *what* is represented with *how* it is represented, leading systems that share identical latent features to be erroneously deemed dissimilar. This work is the first to uncover this mechanism and argues that alignment should be based on recoverable sparse latent features rather than raw, mixed activations. Leveraging compressive sensing theory, random projection analysis, and closed-form derivations, we prove that under sparsity assumptions, the original features can be exactly reconstructed, whereas ignoring the superposition structure yields substantially deflated alignment scores and even incorrect similarity rankings. Our findings establish a new paradigm for evaluating representational similarity in the presence of superposition.
This work addresses a critical limitation in low-rank optimization methods such as GaLore, which rely on the assumption that the gradient subspace drifts slowly and is reproducible. We demonstrate for the first time that this subspace is largely unidentifiable in practice: only approximately 39 out of 128 directional components can be stably reproduced, with the remainder dominated by estimation noise. From a coordinate transformation perspective, we propose LDAdam, an optimizer that correctly propagates Adam’s internal states into newly identified subspaces. The efficacy of LDAdam is substantiated through subspace distance metrics, spectral analysis, and cross-architecture experiments. On a 1B-parameter model, LDAdam achieves a perplexity of 18.7, significantly outperforming the best GaLore configuration (19.3), thereby validating the importance of proper state transfer across subspaces.
This work addresses the lack of a unified theoretical explanation for the alignment phenomenon observed between weight subspaces of adjacent layers in deep neural networks. Building from first principles through geometric invariant theory, the study establishes a unified geometric framework that identifies this alignment structure as a flag manifold and proves that the dimension of subspace intersections is the unique reparameterization-invariant observable. Leveraging Lie bracket analysis and a novel weight-space diagnostic tool, the theory rigorously demonstrates the mathematical necessity of subspace metrics and provides an explanation for the hierarchical structure underlying Neural Collapse. Experimental validation across multilayer perceptrons, residual networks, and pretrained language models confirms the theoretical predictions, with the proposed diagnostic metric accurately capturing internal alignment properties without requiring forward propagation.
This work addresses the high activation memory cost that constitutes a major bottleneck in large-batch training of large language models. Existing compression methods are limited by their neglect of the spectral structure of activations. To overcome this, we propose a principal–random subspace decomposition approach: it preserves critical information in the principal component subspace via singular value decomposition (SVD) while approximating the tail components through unbiased random sampling in the orthogonal complement subspace. Our method establishes, for the first time, a theoretical link between subspace projection and rapid convergence, and introduces an exact scaling factor to minimize the variance of gradient estimates. Experiments demonstrate up to 36% activation memory savings during both pretraining and fine-tuning, with negligible performance degradation and minimal computational overhead.
This work addresses the challenge that existing neural network width theories fail to guarantee the generalization of widening directions identified during training. Focusing on function-preserving residual expansions, the study investigates the alignment between training and test gradients, introducing the notion of “effective alignment dimension” to characterize the signal-to-noise geometric structure of activation gradients. For the first time, this measurable quantity enables a high-probability guarantee of improved test risk under finite samples—without requiring assumptions on covariance spectra or predetermined width growth rates. The theoretical analysis leverages mean-variance decomposition of inner products of activation gradients within a residual expansion framework. Experiments on LLaMA-style Transformers, Pythia, and ResNet-20 demonstrate that wider models exhibit higher effective alignment dimensions and lower empirical misalignment, with this metric accurately predicting both the direction and magnitude of held-out loss changes.