🤖 AI Summary
Deep neural networks (DNNs) suffer from poor interpretability and lack of behavioral controllability. To address this, we propose an integrated “Discover–Verify–Control” framework: first, unsupervised extraction of semantically coherent native feature bases from intermediate layers using an ℓ₁-sparse autoencoder; second, causal verification and feature-level saliency visualization via Grad-FAM, a novel gradient-based attribution method; third, precise and persistent feature editing directly in weight space, accompanied by derivation of a critical suppression threshold α_crit for fine-grained robustness diagnosis. This closed-loop paradigm enables end-to-end, verifiable, and controllable intervention—from feature discovery to causally grounded editing—for the first time. We validate cross-architectural generalizability, causal fidelity, and robustness of the edits on ResNet-18 and ViT-B/16, demonstrating substantial improvements in DNN transparency and controllability.
📝 Abstract
Deep neural networks achieve impressive performance but remain difficult to interpret and control. We present SALVE (Sparse Autoencoder-Latent Vector Editing), a unified "discover, validate, and control" framework that bridges mechanistic interpretability and model editing. Using an $ell_1$-regularized autoencoder, we learn a sparse, model-native feature basis without supervision. We validate these features with Grad-FAM, a feature-level saliency mapping method that visually grounds latent features in input data. Leveraging the autoencoder's structure, we perform precise and permanent weight-space interventions, enabling continuous modulation of both class-defining and cross-class features. We further derive a critical suppression threshold, $α_{crit}$, quantifying each class's reliance on its dominant feature, supporting fine-grained robustness diagnostics. Our approach is validated on both convolutional (ResNet-18) and transformer-based (ViT-B/16) models, demonstrating consistent, interpretable control over their behavior. This work contributes a principled methodology for turning feature discovery into actionable model edits, advancing the development of transparent and controllable AI systems.