Score
Designs and implements methods and analyses that identify, score, and causally test the contribution of individual neurons or neuron pre-activations to model behavior. This includes creating attribution and importance-scoring metrics (e.g., sensitivity, first-order Taylor approximations, variance-based measures), performing neuron ablations or interventions to measure causal output changes, tracking layer-wise or dataset-driven distributional shifts, and combining metrics into target-evidence rankings to prioritize neurons for adaptation or further study.
This work addresses a critical limitation in existing neuron-level concept explanation methods, which often assume that all neurons possess clear functional roles, thereby overlooking redundant or misleading neurons that can distort interpretations of model decision-making. To overcome this, the authors propose the Select-Hypothesize-Verify (SHV) framework: it first selects the most representative samples based on activation distributions, then generates natural language concept hypotheses, and finally validates these hypotheses through a neuron activation verification mechanism. SHV introduces, for the first time, a systematic pipeline for concept validation, effectively identifying and focusing on neurons with genuine semantic meaning. Experimental results demonstrate that concepts produced by SHV activate target neurons at 1.5 times the rate of state-of-the-art methods, substantially improving the accuracy and reliability of model interpretations.
Observational neuroscience studies suffer from persistent confounding, selection bias, and batch effects, undermining causal attribution and generalizability. To address this, we introduce the first end-to-end causal framework specifically designed for neuroscience—spanning experimental design, data acquisition, and modeling analysis. The framework unifies potential outcomes models, causal graphical models, intervention logic, and observational identification techniques, and is rigorously validated using multicenter neuroimaging data. We propose causal-aware experimental design principles and a standardized analytical protocol that substantially improve causal inference credibility and cross-site reproducibility. Moving beyond traditional correlational paradigms, our framework achieves key advances in interpretability, clinical translatability, and methodological robustness. It establishes a foundational paradigm for causal neuroscience research, with direct implications for psychiatry, mental health, and related domains.
Standard optogenetic analyses discard temporal information, thereby limiting causal inference to coarse-grained effects. To address this, we develop a nonparametric causal inference framework that—novel in neuroscience—adapts the “run-length effect” methodology from mobile health. Our approach introduces history-restricted marginal structural models and a taxonomy of identifiable causal effects, unifying treatment of both open-loop static and closed-loop dynamic intervention designs while robustly handling violations of the positivity assumption. The method integrates inverse-probability weighting, doubly-robust estimation at multiple time points, formal hypothesis testing, and computationally efficient implementation. Applied to real neural data, the framework uncovers fine-grained, temporally resolved causal effects of optogenetic interventions on behavior—effects entirely obscured by conventional analyses. It enjoys statistical consistency and asymptotic theoretical guarantees, substantially expanding the scope of scientifically answerable causal questions in optogenetics and systems neuroscience.
Causal interactions in brain networks exhibit inherent nonlinearity and time-varying dynamics, which conventional linear Granger causality methods fail to capture effectively. To address this limitation, we propose TAR4C—a novel framework that integrates threshold autoregressive (TAR) modeling into the Granger causality paradigm for joint modeling and interpretable inference of directional, nonlinear, and time-varying causal relationships in brain networks. Evaluated on multichannel EEG data recorded during motor execution and motor imagery tasks, TAR4C robustly identifies cross-subject consistent, task-discriminative causal connectivity patterns: both conditions rely on primary motor cortex-driven causality, yet motor imagery lacks the sensorimotor feedback-mediated regulatory pathways observed during motor execution. TAR4C thus establishes a new paradigm for dynamic, mechanistically interpretable causal analysis of functional brain networks.
In causal effect estimation, the absence of standardized hyperparameter tuning evaluation criteria impedes reliable model selection and creates a substantial gap between commonly used metrics and true performance. This paper systematically investigates the interplay between hyperparameter tuning and evaluation, jointly analyzing estimators (T-/X-/R-Learner), base learners (random forests, gradient boosting, neural networks), and evaluation metrics (IPW, DR, PEHE) across four benchmark datasets. Key findings are: (1) thorough hyperparameter tuning eliminates performance differences among mainstream causal estimators; (2) the choice of evaluation strategy exerts greater influence on final performance than either the estimator type or base learner architecture; and (3) existing evaluation metrics underestimate the performance gain from optimal model selection by over 35% on average. These results demonstrate that hyperparameter tuning is the primary determinant of causal estimation accuracy, underscoring an urgent need for more robust, theoretically grounded evaluation paradigms in causal machine learning.
Existing approaches struggle to uncover the causal influence of hidden neurons on neural network outputs, as activation patterns alone are insufficient to decipher internal computational mechanisms. This work proposes CODEC, a novel method that—unlike prior activation-based analyses—decouples network behavior into interpretable, sparse contribution modes by integrating contribution decomposition with sparse autoencoders. Applying this framework, the study reveals cross-layer causal computation pathways and demonstrates its efficacy in both image classification and retinal neural activity modeling. The approach enables precise intervention and visualization of intermediate layers, uncovers a progressive decoupling of positive and negative contributions in deeper layers, and elucidates how compositional interactions among intermediate neurons give rise to dynamic receptive fields.
This work addresses the lack of direct validation regarding whether existing neuron attribution methods genuinely identify neurons causally important to model behavior. The authors propose a paired causal auditing framework based on one-shot zero-ablation interventions, integrating contrastive evaluations of harmful versus harmless behaviors with randomized controlled trials to systematically assess attribution efficacy across five large language models. Results demonstrate that the evaluated methods can effectively instill safety-aligned refusal capabilities with low benign rejection rates while preserving linguistic fluency. Notably, the sets of neurons activated by different attribution methods exhibit minimal overlap, suggesting that refusal mechanisms reside within redundant subspaces. Furthermore, high ranking stability does not guarantee strong causal validity, underscoring the necessity of explicit causal verification for attribution techniques.
Existing analyses of neurons in vision-language models are largely confined to single-task settings, overlooking the influence of task-specific attention heads on feedforward neuron writing. This limitation exacerbates neuronal polysemy in multitask scenarios and hinders accurate identification and effective intervention on task-critical neurons. To address this, this work proposes HONES, a framework that jointly models the effects of attention heads and neuron writing for the first time. HONES ranks neurons based on their causal writing contributions under specific tasks and employs lightweight scaling to enable gradient-free, task-aware neuron attribution and control. Experiments across four multimodal tasks and two mainstream architectures demonstrate that HONES more accurately identifies task-critical neurons and significantly improves intervention efficacy.
This study addresses the challenge of distinguishing whether trial-to-trial neuronal variability arises from measurement noise or reflects genuine changes in underlying activation patterns. To this end, the authors propose a two-sample test based on the covariance matrix of functional principal component scores, extending it to paired experimental designs. This approach represents the first application of eigen-decomposition to assess structural consistency in functional data, effectively capturing dynamic trial-level variations that conventional dimensionality reduction methods overlook. Simulations demonstrate superior performance over existing techniques across diverse scenarios. When applied to 157 neural trials, the method significantly detected variability in latent activation patterns that cannot be attributed to sampling noise alone.