Score
Design and implement methods and tools to extract, summarize, visualize, and statistically analyze internal layer- and neuron-level activation signals from neural models — including per-sample activation vectors, activation distributions and sparsity metrics, activation functions, class activation maps, dominant directions and spectral/geometry of activation space, and activation tracing across modules. Use these analyses to probe associations between activations and model outputs or concepts, cluster and flag outlier or backdoor-influenced samples, track signal propagation or amplification/suppression across layers, and support interventions such as quantization, monitoring, or sanitization.
This work addresses a critical limitation in existing neuron-level concept explanation methods, which often assume that all neurons possess clear functional roles, thereby overlooking redundant or misleading neurons that can distort interpretations of model decision-making. To overcome this, the authors propose the Select-Hypothesize-Verify (SHV) framework: it first selects the most representative samples based on activation distributions, then generates natural language concept hypotheses, and finally validates these hypotheses through a neuron activation verification mechanism. SHV introduces, for the first time, a systematic pipeline for concept validation, effectively identifying and focusing on neurons with genuine semantic meaning. Experimental results demonstrate that concepts produced by SHV activate target neurons at 1.5 times the rate of state-of-the-art methods, substantially improving the accuracy and reliability of model interpretations.
This study addresses the challenge of distinguishing whether trial-to-trial neuronal variability arises from measurement noise or reflects genuine changes in underlying activation patterns. To this end, the authors propose a two-sample test based on the covariance matrix of functional principal component scores, extending it to paired experimental designs. This approach represents the first application of eigen-decomposition to assess structural consistency in functional data, effectively capturing dynamic trial-level variations that conventional dimensionality reduction methods overlook. Simulations demonstrate superior performance over existing techniques across diverse scenarios. When applied to 157 neural trials, the method significantly detected variability in latent activation patterns that cannot be attributed to sampling noise alone.
Existing neural network probing methods often rely on input perturbations or parameter analysis, which struggle to uncover structured information embedded in intermediate representations. This work proposes APEX, a novel probing paradigm that perturbs hidden activations during inference while keeping both inputs and model parameters fixed. APEX formalizes activation perturbation as a general probing framework, unifying and extending prior approaches such as input perturbation as special cases, and enabling a controllable transition from sample-dependent to model-dependent behavioral analysis. Experiments demonstrate that APEX effectively quantifies representational structure, distinguishes models trained on structured versus random labels, reveals semantically coherent prediction transitions, and precisely identifies the concentration of predictions toward target classes in backdoor attacks.
This paper addresses the challenge of identifying hierarchical, distributed activation patterns in neural networks. To overcome limitations of neuron-level or hand-crafted interpretable feature analyses, we propose Neural Activation Pattern (NAP) modeling based on full-layer activation distributions. Methodologically, we introduce normalized preprocessing, kernel density estimation for distribution modeling, SNR-adaptive distance metrics, and hierarchical clustering. For the first time, our approach uncovers a continuous activation manifold in neural communication receivers that is dominantly governed by signal-to-noise ratio (SNR), demonstrating its physical interpretability. Experiments show that NAP significantly improves separation between in-distribution and out-of-distribution samples, empirically confirming SNR as a critical implicit factor learned by the model. The method enhances generalization, interpretability, and reliability diagnostics—enabling robust, physics-informed analysis of deep neural representations.
This work addresses the limitation of existing output-confidence–based fault detection methods, which often fail to capture internal errors in neural networks. The authors propose Self-Detecting Neural Networks (SDNN), a novel framework that introduces the concept of “spectral drift” to reveal that erroneous predictions manifest as pronounced multi-scale spectral instabilities in internal activations. Spectral features are extracted via short-time Fourier transform, wavelet decomposition, and statistical moments, and a lightweight detector is trained using curriculum learning to establish an end-to-end learnable internal monitoring mechanism. Evaluated on CIFAR-10, SDNN achieves an AUROC of 79.0 ± 25.3%, outperforming baseline methods such as MaxSoftmax and Energy Score by 25–30 percentage points.
Existing neuron-level interpretability methods are often task-specific, require retraining, or are merely descriptive, making systematic evaluation of Transformer internal robustness challenging. This work proposes SYNAPSE, a framework that enables training-free, cross-architecture, and cross-domain neuron analysis and perturbation for the first time. By extracting [CLS] representations from each layer, SYNAPSE trains lightweight linear probes to rank neurons globally and per class, and applies structured perturbations during inference via forward hooks. Experiments reveal that task-relevant information is encoded by broadly overlapping and functionally redundant neuron subsets, alongside category-asymmetric specialization patterns. Notably, minimal perturbations can drastically alter model predictions, exposing inherent vulnerabilities and offering new insights into the internal mechanisms of Transformers.
This study systematically investigates the impact of Gaussian noise injection—varying by location (before or after activation functions) and type (additive or multiplicative)—on the performance of deep feedforward neural networks. By introducing noise at different layers and incorporating pooling mechanisms, the work reveals that activation functions exhibit a nonlinear noise-filtering effect, and that noise placement critically influences model robustness: injecting additive noise before activation yields higher accuracy and is more effectively suppressed, whereas multiplicative noise has a milder effect when applied after activation. Furthermore, early hidden layers contribute more significantly to performance degradation under post-activation noise injection, while pooling strategies consistently enhance performance across all noise configurations.
Deep learning demonstrates superior performance in EEG-based depression detection, yet its black-box nature hinders clinical adoption. This study presents the first systematic comparison of three families of post-hoc explainability methods—Shapley-value-based (DeepSHAP), gradient-based (Integrated Gradients, GradCAM), and perturbation-based (Occlusion, Permutation Feature Importance)—evaluating their attribution consistency and neurophysiological plausibility on EEG time-series data using the InceptionTime model and subject-wise stratified cross-validation. Results reveal high agreement between gradient- and perturbation-based methods, both highlighting right-hemisphere frontal, temporal, and posterior regions. Although DeepSHAP broadly aligns with prior knowledge of major depressive disorder (MDD), it exhibits markedly distinct attribution patterns, underscoring the critical influence of explanation method choice on interpretability outcomes.
This study addresses the challenge in systems neuroscience that traditional frequentist approaches struggle to effectively adjudicate between highly collinear computational models. To overcome this limitation, the authors propose integrating Bayesian model evidence quantification within the Information Processing Pathway Map (IPPM) framework, shifting model selection from null hypothesis testing to direct comparison of relative evidence among competing models. This work presents the first implementation of Bayesian model comparison in IPPM, substantially enhancing the ability to discriminate between collinear models and enabling robust evidence accumulation across experiments. Applied to reconstructing loudness processing pathways in auditory cortex, the Bayesian approach outperforms conventional frequentist methods in both predictive performance and interpretability, offering a more reliable and theoretically coherent paradigm for selecting neural computational models.