Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of semantic sensitivity in TopK sparse autoencoders, where rare features suffer from geometric selection boundary effects induced by the activity budget k. Through factorial experiments, we elucidate this sensitization mechanism and propose a novel pairwise ranking stabilization algorithm based on active margins to effectively restore feature reliability. Experimental results demonstrate that our approach improves the sensitivity of rare features by 8.83 percentage points while preserving baseline reconstruction quality and feature coverage. This work establishes a new paradigm for sparse coding stability within mechanistic interpretability.
📝 Abstract
Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of TopK selection. The active margin, the distance to the cutoff, predicts feature loss without thresholds. Guided by this margin diagnosis, we introduce pairwise rank stabilization. Our method targets ordering failures at the cutoff and improves rare-feature sensitivity by \(8.83\) percentage points, while keeping reconstruction and alive-feature coverage near the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and boundary geometry for stable interpretability.
Problem

Research questions and friction points this paper is trying to address.

Sparse Autoencoders
TopK Selection
Feature Sensitivity
Interpretability Reliability
Active Budget
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse Autoencoders
TopK Selection
Feature Sensitivity
Active Margin
Pairwise Rank Stabilization