🤖 AI Summary
This study addresses the vulnerability of sparse autoencoder (SAE) features in EEG foundation models to perturbation sensitivity, which frequently leads to semantic misinterpretation. To mitigate this, we propose a multi-level validation ladder framework that systematically evaluates the validity of SAE latent variable semantic interpretations by controlling signal energy, specificity, and raw data correlation, integrated with spectral filtering interventions, statistical normalization, and cross-dataset and cross-backbone validations. Our findings reveal that relying solely on perturbation sensitivity for semantic attribution is inherently misleading, and demonstrate that unnormalized alpha-removal responses introduce systematic bias. Ultimately, this work establishes the core conclusion that sensitivity alone cannot define feature semantics, providing a rigorous evaluation paradigm for the interpretability of representations in computational neuroscience.
📝 Abstract
Sparse autoencoders (SAEs) decompose dense model activations into discrete latents, making individual features easy to interpret--and easy to misinterpret. In EEG foundation models, this creates a tempting inference: if removing alpha-band activity strongly changes a latent's activation, one might conclude that the latent represents alpha activity. Across 27 settings spanning three backbones, three EEG datasets, and three network depths, this interpretation initially appears compelling: alpha removal changes latent firing 7.3 times more than an equal-width sham notch (95% CI [6.2, 8.7], bootstrapped over settings). However, the alpha filter also deletes far more signal than the sham. After normalizing by removed spectral energy, the ratio falls to 0.28 (95% CI [0.22, 0.36]) and exceeds one in none of the 27 settings. Latents selected for their response to alpha removal are, on clean EEG, slightly anti-correlated with relative alpha power (mean r = -0.073), giving no support for a simple alpha-detector reading. Motivated by this failure case, we propose a validation ladder for semantic interpretations of SAE latents: it asks in turn whether a latent responds, whether that response survives controlling for how much signal the intervention removes, whether it is specific rather than broadly fragile, and whether the proposed property is visible on unperturbed data--while separately testing the stronger claim that the latent matters to a task classifier. Perturbation sensitivity alone does not establish what an SAE latent represents.