cka-based layer pruning

Designs and implements methods that compute representation similarity between neural network layers—primarily using centered kernel alignment (CKA)—to quantify redundancy and rank layer pairs by similarity. Uses those similarity rankings to select and remove twin or redundant layers (representation-similarity or layer-similarity pruning), typically implemented to perform pruning decisions efficiently (e.g., in a single forward pass).

cka-basedlayerpruning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.12
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

SGLP: A Similarity Guided Fast Layer Partition Pruning for Compressing Large Deep Models

Oct 14, 2024
YL
Yuqi Li
🏛️ Institute of Computing Technology, Chinese Academy of Sciences | Zhejiang University of Technology | Zhejiang University | Southwest University

Existing layer-pruning methods neglect intrinsic inter-layer dependencies and semantic correlations in deep neural networks, leading to severe knowledge loss. To address this, we propose a representation-driven hierarchical partitioning pruning framework for large model compression. First, we leverage Centered Kernel Alignment (CKA) to measure inter-layer representation similarity and integrate Fisher-optimal segmentation to achieve semantically coherent layer partitioning. Second, we design a fine-grained, structure-aware, zero-fine-tuning importance scoring mechanism—GradNorm—operating within each segment to enable efficient, post-training pruning. Our method significantly outperforms state-of-the-art approaches on both image classification and large language model tasks. Under substantial model compression—up to 50% parameter reduction and 40% FLOPs decrease—it incurs less than 0.5% accuracy degradation, demonstrating strong applicability to edge and resource-constrained deployment scenarios.

Compressing large deep models by removing redundant layers efficientlyMaintaining model accuracy while reducing computational requirements significantlyPreserving essential network characteristics through similarity-guided pruning

Tracing Representation Progression: Analyzing and Enhancing Layer-Wise Similarity

Jun 20, 2024
JJ
Jiachen Jiang
🏛️ Ohio State University

This work investigates the evolution of hidden-layer representation similarity in Transformer models and its implications for training and inference. We observe that inter-layer cosine similarity increases as layer distance decreases and correlates positively with prediction confidence. Building on this, we propose *Alignment Training*: an optimization framework that enforces inter-layer representation alignment to accelerate shallow-layer representation saturation, ensure monotonic improvement in layer-wise accuracy, and inherently reveal the minimal depth required for a given task. Our theoretical analysis is grounded in sample-level similarity metrics and a geometric manifold assumption. Notably, we provide the first formal proof that a single top-layer classifier suffices for multi-exit inference. Empirically, on both vision and NLP benchmarks, our method matches the accuracy of dedicated multi-classifier multi-exit architectures while substantially reducing inference latency and parameter redundancy.

Neural Network InterpretabilitySimilarity EnhancementTransformer Models

Differentiable Optimization of Similarity Scores Between Models and Brains

Jul 09, 2024
NC
Nathan Cloos
🏛️ MIT | NYU | HIH Tübingen

Quantifying and interpreting representational similarity between biological systems (e.g., neural activity) and artificial systems (e.g., deep neural networks) remains challenging due to ambiguities in metric choice, interpretability, and functional relevance. Method: We propose the first end-to-end differentiable optimization framework that directly maximizes representational similarity scores between model and neural representations. Using theoretical analysis and synthetic data inversion, we systematically characterize how CKA, angular Procrustes, and normalized Bures similarity (NBS) differentially weight principal component variance. Contributions: We show CKA strongly biases toward high-variance components, whereas angular Procrustes captures low-variance neural dimensions earlier; high similarity scores do not imply functional equivalence, and no universal threshold exists for neuroscientific interpretation; finally, we delineate the feasible score space and hierarchical constraint strengths under multi-metric joint optimization—establishing a theoretical benchmark and practical guidance for representational similarity assessment.

Comparative AnalysisInformation ProcessingSystem Similarity

Evaluating Representational Similarity Measures from the Lens of Functional Correspondence

Nov 21, 2024
YB
Yiqing Bo
🏛️ UC San Diego | University of Pennsylvania

This study addresses the challenge of comparing high-dimensional neural representations across neuroscience and artificial intelligence: specifically, how to select similarity measures that best reveal functional correspondences and divergences. We systematically evaluate eight mainstream representational similarity metrics—including linear CKA, Procrustes distance, CCA, inner-product kernel, and nearest-neighbor alignment—against behavioral functional alignment (e.g., recognition accuracy, generalization, robustness) as a ground-truth benchmark. Our evaluation spans both biological neural data and artificial neural network models. Results show that geometry-sensitive metrics—particularly linear CKA and Procrustes distance—consistently outperform predictive metrics, achieving superior alignment with human behavioral performance and effectively distinguishing trained versus untrained models. In contrast, linear predictivity exhibits only moderate behavioral alignment. This work establishes the first behavior-driven representational similarity benchmark, providing a principled, cross-domain methodology for mechanistic interpretation and comparative analysis of neural computation.

Assessing alignment between similarity measures and behavioral outcomesEvaluating representational similarity metrics for neural data comparisonIdentifying metrics that best differentiate model behaviors and training

Neural representational similarity measurement suffers from inconsistent nomenclature and implementation, severely impeding cross-study comparability and result reproducibility. To address this, we propose the first dynamically evolving naming standard framework that enables scalable, verifiable, and globally unique identifiers for similarity measures—overcoming the rigidity of static standards in rapidly advancing domains. We develop an open-source Python benchmarking platform integrating 14 mainstream toolkits (encompassing ~100 distinct measures) and provide unified formal modeling and explicit differentiation among 12+ variants of key methods (e.g., CKA). Leveraging a modular architecture and multi-package-compatible interfaces, our platform supports out-of-the-box standardized computation and evaluation. This work significantly enhances transparency, reproducibility, and efficiency of cross-study comparisons in representational similarity analysis, establishing foundational infrastructure for neural representation research.

Addressing naming and implementation inconsistencies across studiesProviding a framework for benchmarking and comparing similarity methodsStandardizing diverse similarity measures in evolving fields

Latest Papers

What's happening recently
View more

This work proposes the “similarity triangle” framework to address the limitations of existing neural representation comparison methods, which often adopt a narrow perspective and fail to comprehensively capture the similarity of internal model mechanisms. The framework unifies three complementary dimensions—static (e.g., CKA or Procrustes), functional (e.g., linear probing connectivity and prediction similarity), and sparsity-based (pruning-induced robustness)—to systematically evaluate the representational geometry of CNNs, Vision Transformers (ViTs), and vision-language models. Experiments reveal that model architecture predominantly governs representational similarity, yielding distinct architectural family clusters on ImageNetV2 and CIFAR-10. Furthermore, CKA self-similarity strongly correlates with task performance, and pruning not only uncovers shared computational cores across models but also exerts a representational regularization effect.

internal mechanismsmodel comparisonneural network representations

Manifold Approximation leads to Robust Kernel Alignment

Oct 26, 2025
MT
Mohammad Tariqul Islam
🏛️ MIT

Existing Centered Kernel Alignment (CKA) methods neglect the underlying manifold structure of data and rely on heuristic designs, resulting in poor stability across data scales and limited interpretability. To address this, we propose Manifold-aware Kernel Alignment (MKA), the first framework to systematically integrate manifold geometry into kernel alignment: it constructs scale-adaptive kernels via local manifold approximation and establishes a theoretically grounded, geometrically consistent alignment criterion. MKA requires no hyperparameter tuning and significantly improves robustness and discriminability in representation comparison—especially under few-shot and multi-scale regimes—on both synthetic and real-world benchmarks. Experiments demonstrate its superiority in tasks such as representation equivalence testing and neural representational analysis, delivering more reliable and interpretable similarity metrics. By bridging manifold learning with kernel-based representation comparison, MKA provides a principled tool for representation learning and computational neuroscience.

Addressing CKA's sensitivity to data scale and heuristicsImproving kernel alignment by incorporating manifold geometryProviding robust foundation for measuring representations

This work proposes a joint optimization pruning strategy that moves beyond the conventional paradigm of assessing weight importance solely based on magnitude. Recognizing that performance degradation caused by weight removal can be partially compensated through adjustments to adjacent biases, the method simultaneously prunes weights and computes optimal bias perturbations via automatic differentiation to minimize accuracy loss. By explicitly accounting for the interplay between weights and biases, this approach provides a more accurate measure of weight significance. Extensive experiments demonstrate that the proposed technique consistently outperforms state-of-the-art pruning methods across various models and tasks, achieving superior accuracy and robustness—particularly under high pruning ratios.

bias compensationneural networkspruning

Existing studies lack a quantitative, discriminability-oriented comparison of representational similarity measures (RSMs) across diverse model families (e.g., CNNs, Transformers, biologically inspired models). Method: We propose a unified discriminability evaluation framework grounded in signal detection theory (d′), silhouette coefficient, and ROC-AUC, integrating major RSMs—including Representational Similarity Analysis (RSA), linear predictivity, Procrustes alignment, and soft matching. Contribution/Results: Our analysis reveals, for the first time, a positive correlation between an RSM’s alignment constraint strength and its ability to separate models by architecture or training paradigm: soft matching achieves top discriminability, while non-fitting methods (e.g., RSA) also exhibit strong separation performance. The framework provides interpretable, reproducible criteria for selecting RSMs in cross-model comparisons and model–brain alignment studies.

Assessing separability capacity of various similarity measures using quantitative frameworkEvaluating discriminative power of representational similarity metrics across model familiesSystematically comparing metric performance across architectures and training regimes

Accuracy-Preserving CNN Pruning Method under Limited Data Availability

Nov 13, 2025
DY
Daisuke Yasui
🏛️ National Defense Academy of Japan

In data-scarce scenarios, conventional LRP-based CNN channel pruning suffers from substantial accuracy degradation and limited pruning ratios. To address this, we propose a fine-tuning-free, interpretability-driven pruning framework. Our method dynamically refines per-layer channel relevance scores by jointly leveraging structural priors from pre-trained models and an LRP-based importance re-evaluation mechanism—without requiring additional labeled data. This yields more robust channel importance estimation under low-data conditions. Experiments on ImageNet subsets and CIFAR benchmarks demonstrate that our approach achieves, on average, a 23.6% higher pruning ratio and reduces accuracy loss by 58.4% compared to state-of-the-art LRP-based pruning methods. It thus significantly overcomes the performance bottleneck of traditional LRP pruning in few-shot settings, delivering a practical, high-accuracy, high-compression solution suitable for resource-constrained edge deployment.

Achieving higher pruning rates while preserving model performanceAddressing accuracy degradation in existing LRP-based pruning approachesDeveloping CNN pruning method that maintains accuracy with limited data

Hot Scholars

JZ

Jun Zhang

ByteDance
Speech RecognitionAcoustic Event DetectionBCI
XZ

Xueyao Zhang

The Chinese University of Hong Kong, Shenzhen
Deep LearningSpeechSinging VoiceMusic
MK

Mykhailo Koshil

Ph.D. Candidate at "AutoML for Science" group | University of Tübingen
Tabular DataFoundation ModelsAutoML
SL

Sijia Li

Institute of Information Engineering, Chinese Academy of Sciences
ZW

Zhizheng Wu

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing