Score
Designs and executes analyses that compute centered kernel alignment in its linear form (linear CKA) to quantify similarity or alignment between vector representations—across layers, seeds, checkpoints, or individual channels. Uses linear CKA as a probing tool and builds pipelines that relate alignment values to downstream metrics (e.g., predict accuracy), track alignment changes across initialization or training sweeps, and diagnose channel‑level drift or closure via patterns in weight or activation alignment.
Neural representational similarity measurement suffers from inconsistent nomenclature and implementation, severely impeding cross-study comparability and result reproducibility. To address this, we propose the first dynamically evolving naming standard framework that enables scalable, verifiable, and globally unique identifiers for similarity measures—overcoming the rigidity of static standards in rapidly advancing domains. We develop an open-source Python benchmarking platform integrating 14 mainstream toolkits (encompassing ~100 distinct measures) and provide unified formal modeling and explicit differentiation among 12+ variants of key methods (e.g., CKA). Leveraging a modular architecture and multi-package-compatible interfaces, our platform supports out-of-the-box standardized computation and evaluation. This work significantly enhances transparency, reproducibility, and efficiency of cross-study comparisons in representational similarity analysis, establishing foundational infrastructure for neural representation research.
This work investigates the evolution of hidden-layer representation similarity in Transformer models and its implications for training and inference. We observe that inter-layer cosine similarity increases as layer distance decreases and correlates positively with prediction confidence. Building on this, we propose *Alignment Training*: an optimization framework that enforces inter-layer representation alignment to accelerate shallow-layer representation saturation, ensure monotonic improvement in layer-wise accuracy, and inherently reveal the minimal depth required for a given task. Our theoretical analysis is grounded in sample-level similarity metrics and a geometric manifold assumption. Notably, we provide the first formal proof that a single top-layer classifier suffices for multi-exit inference. Empirically, on both vision and NLP benchmarks, our method matches the accuracy of dedicated multi-classifier multi-exit architectures while substantially reducing inference latency and parameter redundancy.
This study addresses the challenge of comparing high-dimensional neural representations across neuroscience and artificial intelligence: specifically, how to select similarity measures that best reveal functional correspondences and divergences. We systematically evaluate eight mainstream representational similarity metrics—including linear CKA, Procrustes distance, CCA, inner-product kernel, and nearest-neighbor alignment—against behavioral functional alignment (e.g., recognition accuracy, generalization, robustness) as a ground-truth benchmark. Our evaluation spans both biological neural data and artificial neural network models. Results show that geometry-sensitive metrics—particularly linear CKA and Procrustes distance—consistently outperform predictive metrics, achieving superior alignment with human behavioral performance and effectively distinguishing trained versus untrained models. In contrast, linear predictivity exhibits only moderate behavioral alignment. This work establishes the first behavior-driven representational similarity benchmark, providing a principled, cross-domain methodology for mechanistic interpretation and comparative analysis of neural computation.
This paper investigates the underlying causes of linear mode connectivity (LMC) and its relationship with weight matching (WM). Addressing two key questions—whether WM achieves LMC solely by reducing the $L^2$ distance between models, and how it ensures low-loss interpolation paths—the work provides the first analysis from the perspective of singular vector alignment. It reveals that WM preserves singular values while aligning the directions of dominant left and right singular vectors, thereby achieving structural alignment between functionally equivalent models. This mechanism is provably equivalent to activation matching (AM) and outperforms data-dependent straight-through estimators (STE). Empirically, the study demonstrates that WM induces LMC without requiring $L^2$-distance compression. Moreover, it unifies theoretical explanations for model merging and SGD generalization. The findings establish a geometric foundation for WM-based model interpolation and offer new insights into the loss landscape geometry of deep neural networks.
This study investigates the differential contributions of individual layers in large language models (LLMs) during supervised fine-tuning for alignment. Method: We propose Importance-aware Layer Adaptation (ILA), a binary mask learning framework to quantitatively assess layer-wise importance. Contribution/Results: Through systematic analysis, we find—contrary to common assumptions—that alignment primarily reshapes representation styles rather than modifying underlying knowledge. Crucially, key layers identified by ILA exhibit 90% overlap across diverse datasets, demonstrating strong generalizability. Empirically, freezing non-critical layers improves final alignment performance while substantially reducing GPU memory consumption and computational cost. Moreover, fine-tuning only the ILA-identified critical layers achieves 98% of the performance attained by full-model fine-tuning. Our work establishes a new paradigm for efficient, interpretable, and layer-aware LLM alignment.
This work addresses the challenge of ineffective alignment signals in existing CLIP and DINOv2 methods when trained on web-scale noisy datasets like CC12M, where fixed loss weights hinder visual representation learning. To overcome this, the authors propose KALE—the first adaptive alignment mechanism tailored for large-scale noisy data—which dynamically adjusts the CLIP–DINOv2 alignment loss weight to preserve effective gradient signals without requiring dataset-specific hyperparameter tuning. Integrating a kernel alignment framework, an adaptive loss balancing controller, and an aggressive learning rate decay strategy, KALE achieves a +2.00 average improvement in zero-shot performance across 11 standard benchmarks on a 3.3M-image subset, substantially outperforming KUEA (+1.29) and delivering consistent gains in SVHN linear probing and image–text retrieval tasks.
Do different large language model architectures encode high-level concepts in a structurally compatible manner? Existing methods struggle to disentangle generic representational similarity from concept-specific alignment. This work proposes Contrastive Differential Centered Kernel Alignment (CKA_Delta), a training-free approach that isolates concept-specific structural convergence signals by leveraging sample-level contrastive differences. Applying CKA_Delta reveals, for the first time, a decoupling between geometric convergence and functional transfer: across six conceptual domains, models exhibit moderate geometric convergence alongside near-perfect functional transfer. CKA_Delta substantially outperforms standard CKA in identifying cross-architecture concept alignment and successfully flags outlier models such as Gemma (AUC = 0.79), offering a novel diagnostic tool for model analysis.
This work investigates how the alignment between eigenvectors of kernel matrices and learning targets influences generalization performance in finite-sample settings within kernel methods. Framed in kernel ridge regression, the study leverages matrix perturbation theory and spectral analysis to establish, for the first time, finite-sample generalization error bounds that jointly account for feature alignment, eigenvalue magnitude, and spectral gap. The theoretical analysis reveals that strong generalization arises when there is high alignment, large eigenvalues, or a pronounced spectral gap. Notably, under high-rank kernels, low reconstruction error proves insufficient for predicting generalization capability, thereby exposing a fundamental limitation of conventional reconstruction-based metrics.
Existing activation alignment methods struggle to capture differences in the sensitivity of neural representations to local stimulus perturbations and thus fail to reflect how systems leverage local evidence for discrimination. This work proposes a novel analytical framework based on locally decodable information, integrating Fisher information, pullback metrics, and log-spectral distances on the SPD manifold to construct the Spectral Riemannian Alignment Score (S-RAS). S-RAS provides, for the first time, a minimal, dataset-level summary of neural representational sensitivity from the perspective of local discriminative tasks, with guaranteed multiplicative consistency. The method successfully aligns corresponding layers across independently trained networks, enables transferable class-conditional probing, reveals representational differences between standard and robustly trained models, and uncovers stimulus coordinate family effects in mouse visual cortex.
本文通过学习核对齐来解决多类贝叶斯分类中预先选择核的问题,使用协作学习和推理框架,并引入马氏距离改进了分类性能。