ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of aligning sparse autoencoder (SAE) features with target concepts, which typically relies on post-hoc probing. To overcome this limitation, we propose introducing a differentiable feature-category affinity matrix during training to automatically assign categories and guide the encoder toward learning class-separable representations. Methodologically, our approach integrates a differentiable Top-k operator with CLIP ViT-L/14 embeddings, enabling gradients to flow through the feature selection process for end-to-end joint optimization. This yields a category-annotated dictionary without requiring post-hoc probing. Experimental evaluations on ImageNet demonstrate the accuracy of the affinity matrix, support direct classification prediction, and reveal significantly improved class separability under probe perturbation assessments.
📝 Abstract
Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mitigation, and concept editing. A central challenge for these editing and steering methods is reliably matching features to target concepts; most current approaches address this by computing post-hoc scores over an already-trained, frozen dictionary. We instead introduce ClasSAE, a novel method that both automatically assigns classes to features and guides the encoder toward class-separable representations during training. Specifically, we apply a differentiable top-$k$ operator to a trainable feature--class affinity matrix with per-feature budgets, coupling the features selected for each sample to the classes they are trained to represent. Because gradients flow through the selection of active features rather than only through their magnitudes, the encoder and the affinity matrix co-adapt rather than being fit in separate stages. The result is a dictionary that is both class-separable and class-annotated, with no need for post-hoc probing. We propose three variants for enforcing sparsity within this framework, which achieve comparable overall performance with slightly different trade-offs. Using CLIP ViT-L/14 embeddings on ImageNet, we show that the learned affinity matrix agrees closely with an independently estimated post-hoc feature-class matrix computed on held-out data. The model also supports direct class prediction from the encoder and affinity matrix alone, without a separately fitted classifier, and its more class-aligned encoder yields improved separation in Targeted Probe Perturbation evaluations. https://github.com/St0pien/ClasSAE.
Problem

Research questions and friction points this paper is trying to address.

Sparse Autoencoders
Feature-Class Matching
Concept Editing
Interpretability
Class-Separable Representations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
Differentiable Top-k
Feature-Class Affinity
Class-Aligned Representations
Mechanistic Interpretability