linear cka probing

Designs and executes analyses that compute centered kernel alignment in its linear form (linear CKA) to quantify similarity or alignment between vector representations—across layers, seeds, checkpoints, or individual channels. Uses linear CKA as a probing tool and builds pipelines that relate alignment values to downstream metrics (e.g., predict accuracy), track alignment changes across initialization or training sweeps, and diagnose channel‑level drift or closure via patterns in weight or activation alignment.

linearckaprobing

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Neural representational similarity measurement suffers from inconsistent nomenclature and implementation, severely impeding cross-study comparability and result reproducibility. To address this, we propose the first dynamically evolving naming standard framework that enables scalable, verifiable, and globally unique identifiers for similarity measures—overcoming the rigidity of static standards in rapidly advancing domains. We develop an open-source Python benchmarking platform integrating 14 mainstream toolkits (encompassing ~100 distinct measures) and provide unified formal modeling and explicit differentiation among 12+ variants of key methods (e.g., CKA). Leveraging a modular architecture and multi-package-compatible interfaces, our platform supports out-of-the-box standardized computation and evaluation. This work significantly enhances transparency, reproducibility, and efficiency of cross-study comparisons in representational similarity analysis, establishing foundational infrastructure for neural representation research.

Addressing naming and implementation inconsistencies across studiesProviding a framework for benchmarking and comparing similarity methodsStandardizing diverse similarity measures in evolving fields

Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity

Jun 20, 2024
JJ
Jiachen Jiang
🏛️ Ohio State University

This work investigates the evolution of hidden-layer representation similarity in Transformer models and its implications for training and inference. We observe that inter-layer cosine similarity increases as layer distance decreases and correlates positively with prediction confidence. Building on this, we propose *Alignment Training*: an optimization framework that enforces inter-layer representation alignment to accelerate shallow-layer representation saturation, ensure monotonic improvement in layer-wise accuracy, and inherently reveal the minimal depth required for a given task. Our theoretical analysis is grounded in sample-level similarity metrics and a geometric manifold assumption. Notably, we provide the first formal proof that a single top-layer classifier suffices for multi-exit inference. Empirically, on both vision and NLP benchmarks, our method matches the accuracy of dedicated multi-classifier multi-exit architectures while substantially reducing inference latency and parameter redundancy.

Neural Network InterpretabilitySimilarity EnhancementTransformer Models

Evaluating Representational Similarity Measures from the Lens of Functional Correspondence

Nov 21, 2024
YB
Yiqing Bo
🏛️ UC San Diego | University of Pennsylvania

This study addresses the challenge of comparing high-dimensional neural representations across neuroscience and artificial intelligence: specifically, how to select similarity measures that best reveal functional correspondences and divergences. We systematically evaluate eight mainstream representational similarity metrics—including linear CKA, Procrustes distance, CCA, inner-product kernel, and nearest-neighbor alignment—against behavioral functional alignment (e.g., recognition accuracy, generalization, robustness) as a ground-truth benchmark. Our evaluation spans both biological neural data and artificial neural network models. Results show that geometry-sensitive metrics—particularly linear CKA and Procrustes distance—consistently outperform predictive metrics, achieving superior alignment with human behavioral performance and effectively distinguishing trained versus untrained models. In contrast, linear predictivity exhibits only moderate behavioral alignment. This work establishes the first behavior-driven representational similarity benchmark, providing a principled, cross-domain methodology for mechanistic interpretation and comparative analysis of neural computation.

Assessing alignment between similarity measures and behavioral outcomesEvaluating representational similarity metrics for neural data comparisonIdentifying metrics that best differentiate model behaviors and training

Analysis of Linear Mode Connectivity via Permutation-Based Weight Matching

Feb 06, 2024
AI
Akira Ito
🏛️ Nippon Telegraph and Telephone Corporation

This paper investigates the underlying causes of linear mode connectivity (LMC) and its relationship with weight matching (WM). Addressing two key questions—whether WM achieves LMC solely by reducing the $L^2$ distance between models, and how it ensures low-loss interpolation paths—the work provides the first analysis from the perspective of singular vector alignment. It reveals that WM preserves singular values while aligning the directions of dominant left and right singular vectors, thereby achieving structural alignment between functionally equivalent models. This mechanism is provably equivalent to activation matching (AM) and outperforms data-dependent straight-through estimators (STE). Empirically, the study demonstrates that WM induces LMC without requiring $L^2$-distance compression. Moreover, it unifies theoretical explanations for model merging and SGD generalization. The findings establish a geometric foundation for WM-based model interpolation and offer new insights into the loss landscape geometry of deep neural networks.

Analyzes linear mode connectivity via weight matching permutationsCompares weight matching with dataset-dependent permutation methodsExplores singular vector alignment for model functionality retention

Understanding Layer Significance in LLM Alignment

Oct 23, 2024
GS
Guangyuan Shi
🏛️ The Hong Kong Polytechnic University | Du Xiaoman Financial

This study investigates the differential contributions of individual layers in large language models (LLMs) during supervised fine-tuning for alignment. Method: We propose Importance-aware Layer Adaptation (ILA), a binary mask learning framework to quantitatively assess layer-wise importance. Contribution/Results: Through systematic analysis, we find—contrary to common assumptions—that alignment primarily reshapes representation styles rather than modifying underlying knowledge. Crucially, key layers identified by ILA exhibit 90% overlap across diverse datasets, demonstrating strong generalizability. Empirically, freezing non-critical layers improves final alignment performance while substantially reducing GPU memory consumption and computational cost. Moreover, fine-tuning only the ILA-identified critical layers achieves 98% of the performance attained by full-model fine-tuning. Our work establishes a new paradigm for efficient, interpretable, and layer-aware LLM alignment.

Identify critical layers in LLM alignment processImprove fine-tuning efficiency by focusing on key layersUnderstand alignment impact on model behavior vs knowledge

Latest Papers

What's happening recently
View more

This work addresses the challenge of ineffective alignment signals in existing CLIP and DINOv2 methods when trained on web-scale noisy datasets like CC12M, where fixed loss weights hinder visual representation learning. To overcome this, the authors propose KALE—the first adaptive alignment mechanism tailored for large-scale noisy data—which dynamically adjusts the CLIP–DINOv2 alignment loss weight to preserve effective gradient signals without requiring dataset-specific hyperparameter tuning. Integrating a kernel alignment framework, an adaptive loss balancing controller, and an aggressive learning rate decay strategy, KALE achieves a +2.00 average improvement in zero-shot performance across 11 standard benchmarks on a 3.3M-image subset, substantially outperforming KUEA (+1.29) and delivering consistent gains in SVHN linear probing and image–text retrieval tasks.

CLIP-DINOv2 alignmentkernel alignmentloss equilibration

Do different large language model architectures encode high-level concepts in a structurally compatible manner? Existing methods struggle to disentangle generic representational similarity from concept-specific alignment. This work proposes Contrastive Differential Centered Kernel Alignment (CKA_Delta), a training-free approach that isolates concept-specific structural convergence signals by leveraging sample-level contrastive differences. Applying CKA_Delta reveals, for the first time, a decoupling between geometric convergence and functional transfer: across six conceptual domains, models exhibit moderate geometric convergence alongside near-perfect functional transfer. CKA_Delta substantially outperforms standard CKA in identifying cross-architecture concept alignment and successfully flags outlier models such as Gemma (AUC = 0.79), offering a novel diagnostic tool for model analysis.

concept encodingcross-architecture comparisongeometric-functional universality

This work investigates how the alignment between eigenvectors of kernel matrices and learning targets influences generalization performance in finite-sample settings within kernel methods. Framed in kernel ridge regression, the study leverages matrix perturbation theory and spectral analysis to establish, for the first time, finite-sample generalization error bounds that jointly account for feature alignment, eigenvalue magnitude, and spectral gap. The theoretical analysis reveals that strong generalization arises when there is high alignment, large eigenvalues, or a pronounced spectral gap. Notably, under high-rank kernels, low reconstruction error proves insufficient for predicting generalization capability, thereby exposing a fundamental limitation of conventional reconstruction-based metrics.

eigen-alignmentfinite-sample analysisgeneralization error

Existing activation alignment methods struggle to capture differences in the sensitivity of neural representations to local stimulus perturbations and thus fail to reflect how systems leverage local evidence for discrimination. This work proposes a novel analytical framework based on locally decodable information, integrating Fisher information, pullback metrics, and log-spectral distances on the SPD manifold to construct the Spectral Riemannian Alignment Score (S-RAS). S-RAS provides, for the first time, a minimal, dataset-level summary of neural representational sensitivity from the perspective of local discriminative tasks, with guaranteed multiplicative consistency. The method successfully aligns corresponding layers across independently trained networks, enables transferable class-conditional probing, reveals representational differences between standard and robustly trained models, and uncovers stimulus coordinate family effects in mouse visual cortex.

activation alignmentFisher informationlocal sensitivity

Hot Scholars

ZX

Zhenhua Xu

Zhejiang University
AI SecurityReinforce Learning
AP

Ankita Pasad

NVIDIA
speech and language processingmachine learning
DW

Dongxia Wu

Stanford University
Uncertainty QuantificationAI for ScienceTrustworthy AISpatiotemporal Modeling
DK

Dhruv Kumar

Faculty @ BITS Pilani. Adjunct Faculty @ IIIT Delhi, Ex-Microsoft, Google. PhD @ UMinnesota-TC, USA
Large Language ModelsGenerative AIComputing EducationICT4D
RW

Rui Wang

Institute of Automation, Chinese Academy of Sciences
Biomimetic robotunderwater robotintelligent controlmechatronics