neuron attribution

Designs and implements methods and analyses that identify, score, and causally test the contribution of individual neurons or neuron pre-activations to model behavior. This includes creating attribution and importance-scoring metrics (e.g., sensitivity, first-order Taylor approximations, variance-based measures), performing neuron ablations or interventions to measure causal output changes, tracking layer-wise or dataset-driven distributional shifts, and combining metrics into target-evidence rankings to prioritize neurons for adaptation or further study.

neuronattribution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.27
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$226K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical limitation in existing neuron-level concept explanation methods, which often assume that all neurons possess clear functional roles, thereby overlooking redundant or misleading neurons that can distort interpretations of model decision-making. To overcome this, the authors propose the Select-Hypothesize-Verify (SHV) framework: it first selects the most representative samples based on activation distributions, then generates natural language concept hypotheses, and finally validates these hypotheses through a neuron activation verification mechanism. SHV introduces, for the first time, a systematic pipeline for concept validation, effectively identifying and focusing on neurons with genuine semantic meaning. Experimental results demonstrate that concepts produced by SHV activate target neurons at 1.5 times the rate of state-of-the-art methods, substantially improving the accuracy and reliability of model interpretations.

concept verificationmisleading neuron conceptsneural network interpretability

How causal perspectives can inform problems in computational neuroscience

Mar 12, 2025
EW
Eric W. Bridgeford
🏛️ Stanford University | Johns Hopkins University

Observational neuroscience studies suffer from persistent confounding, selection bias, and batch effects, undermining causal attribution and generalizability. To address this, we introduce the first end-to-end causal framework specifically designed for neuroscience—spanning experimental design, data acquisition, and modeling analysis. The framework unifies potential outcomes models, causal graphical models, intervention logic, and observational identification techniques, and is rigorously validated using multicenter neuroimaging data. We propose causal-aware experimental design principles and a standardized analytical protocol that substantially improve causal inference credibility and cross-site reproducibility. Moving beyond traditional correlational paradigms, our framework achieves key advances in interpretability, clinical translatability, and methodological robustness. It establishes a foundational paradigm for causal neuroscience research, with direct implications for psychiatry, mental health, and related domains.

Addressing causality challenges in observational neuroscience studiesIncorporating causal inference frameworks to handle confounding and biasesProviding practical tools for causal interpretation in neuroscience analysis

Nonparametric causal inference for optogenetics: sequential excursion effects for dynamic regimes

May 28, 2024
GL
Gabriel Loewinger
🏛️ National Institute of Mental Health | NIH | Carnegie Mellon University

Standard optogenetic analyses discard temporal information, thereby limiting causal inference to coarse-grained effects. To address this, we develop a nonparametric causal inference framework that—novel in neuroscience—adapts the “run-length effect” methodology from mobile health. Our approach introduces history-restricted marginal structural models and a taxonomy of identifiable causal effects, unifying treatment of both open-loop static and closed-loop dynamic intervention designs while robustly handling violations of the positivity assumption. The method integrates inverse-probability weighting, doubly-robust estimation at multiple time points, formal hypothesis testing, and computationally efficient implementation. Applied to real neural data, the framework uncovers fine-grained, temporally resolved causal effects of optogenetic interventions on behavior—effects entirely obscured by conventional analyses. It enjoys statistical consistency and asymptotic theoretical guarantees, substantially expanding the scope of scientifically answerable causal questions in optogenetics and systems neuroscience.

Addresses positivity violations in closed-loop optogenetics designsDevelops nonparametric causal inference for optogenetics behavioral experimentsExtends excursion effects to handle dynamic treatment regimes

Nonlinear Causality in Brain Networks: With Application to Motor Imagery vs Execution

Sep 16, 2024
SA
Sipan Aslan
🏛️ King Abdullah University of Science and Technology

Causal interactions in brain networks exhibit inherent nonlinearity and time-varying dynamics, which conventional linear Granger causality methods fail to capture effectively. To address this limitation, we propose TAR4C—a novel framework that integrates threshold autoregressive (TAR) modeling into the Granger causality paradigm for joint modeling and interpretable inference of directional, nonlinear, and time-varying causal relationships in brain networks. Evaluated on multichannel EEG data recorded during motor execution and motor imagery tasks, TAR4C robustly identifies cross-subject consistent, task-discriminative causal connectivity patterns: both conditions rely on primary motor cortex-driven causality, yet motor imagery lacks the sensorimotor feedback-mediated regulatory pathways observed during motor execution. TAR4C thus establishes a new paradigm for dynamic, mechanistically interpretable causal analysis of functional brain networks.

Applying method to EEG data from motor imagery/execution experimentsIdentifying nonlinear causal interactions in time series networksProposing threshold autoregressive modeling for causality detection

Hyperparameter Tuning and Model Evaluation in Causal Effect Estimation

Mar 02, 2023
DM
Damian Machlanski
🏛️ University of Essex

In causal effect estimation, the absence of standardized hyperparameter tuning evaluation criteria impedes reliable model selection and creates a substantial gap between commonly used metrics and true performance. This paper systematically investigates the interplay between hyperparameter tuning and evaluation, jointly analyzing estimators (T-/X-/R-Learner), base learners (random forests, gradient boosting, neural networks), and evaluation metrics (IPW, DR, PEHE) across four benchmark datasets. Key findings are: (1) thorough hyperparameter tuning eliminates performance differences among mainstream causal estimators; (2) the choice of evaluation strategy exerts greater influence on final performance than either the estimator type or base learner architecture; and (3) existing evaluation metrics underestimate the performance gain from optimal model selection by over 35% on average. These results demonstrate that hyperparameter tuning is the primary determinant of causal estimation accuracy, underscoring an urgent need for more robust, theoretically grounded evaluation paradigms in causal machine learning.

Complex model selection involving multiple components complicates causal inferenceLack of consensus on tuning metrics for causal effect estimation modelsNo ideal metric exists for hyperparameter tuning of causal estimators

Latest Papers

What's happening recently
View more

Existing approaches struggle to uncover the causal influence of hidden neurons on neural network outputs, as activation patterns alone are insufficient to decipher internal computational mechanisms. This work proposes CODEC, a novel method that—unlike prior activation-based analyses—decouples network behavior into interpretable, sparse contribution modes by integrating contribution decomposition with sparse autoencoders. Applying this framework, the study reveals cross-layer causal computation pathways and demonstrates its efficacy in both image classification and retinal neural activity modeling. The approach enables precise intervention and visualization of intermediate layers, uncovers a progressive decoupling of positive and negative contributions in deeper layers, and elucidates how compositional interactions among intermediate neurons give rise to dynamic receptive fields.

causal interpretationcontribution decompositionhidden neuron contributions

This work addresses the lack of direct validation regarding whether existing neuron attribution methods genuinely identify neurons causally important to model behavior. The authors propose a paired causal auditing framework based on one-shot zero-ablation interventions, integrating contrastive evaluations of harmful versus harmless behaviors with randomized controlled trials to systematically assess attribution efficacy across five large language models. Results demonstrate that the evaluated methods can effectively instill safety-aligned refusal capabilities with low benign rejection rates while preserving linguistic fluency. Notably, the sets of neurons activated by different attribution methods exhibit minimal overlap, suggesting that refusal mechanisms reside within redundant subspaces. Furthermore, high ranking stability does not guarantee strong causal validity, underscoring the necessity of explicit causal verification for attribution techniques.

attribution scorescausal auditfaithfulness

Existing analyses of neurons in vision-language models are largely confined to single-task settings, overlooking the influence of task-specific attention heads on feedforward neuron writing. This limitation exacerbates neuronal polysemy in multitask scenarios and hinders accurate identification and effective intervention on task-critical neurons. To address this, this work proposes HONES, a framework that jointly models the effects of attention heads and neuron writing for the first time. HONES ranks neurons based on their causal writing contributions under specific tasks and employs lightweight scaling to enable gradient-free, task-aware neuron attribution and control. Experiments across four multimodal tasks and two mainstream architectures demonstrate that HONES more accurately identifies task-critical neurons and significantly improves intervention efficacy.

causal write-in effectsmulti-task vision-language modelsneuron attribution

This study addresses the challenge of distinguishing whether trial-to-trial neuronal variability arises from measurement noise or reflects genuine changes in underlying activation patterns. To this end, the authors propose a two-sample test based on the covariance matrix of functional principal component scores, extending it to paired experimental designs. This approach represents the first application of eigen-decomposition to assess structural consistency in functional data, effectively capturing dynamic trial-level variations that conventional dimensionality reduction methods overlook. Simulations demonstrate superior performance over existing techniques across diverse scenarios. When applied to 157 neural trials, the method significantly detected variability in latent activation patterns that cannot be attributed to sampling noise alone.

eigendecompositionfunctional datalatent activation patterns

Hot Scholars

TS

Tat-Seng Chua

National University of Singapore
Multimedia Information RetrievalLive Social Media Analysis
FS

Fei Shen

National University of Singapore
Controllable GenerationMultimodal Safety
BS

Berrak Sisman

Assistant Professor (ECE & DSAI), Johns Hopkins University
Machine LearningAffective ComputingSpeech SynthesisVoice Conversion
BS

Björn Schuller

Professor, Technische Universität München (TUM) / Imperial College London & CSO, audEERING
Health InformaticsDigital HealthAIAffective Computing
XZ

Xiutian Zhao

Johns Hopkins University
Speech and Language Processing