SALVE: Sparse Autoencoder-Latent Vector Editing for Mechanistic Control of Neural Networks

📅 2025-12-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Deep neural networks (DNNs) suffer from poor interpretability and lack of behavioral controllability. To address this, we propose an integrated “Discover–Verify–Control” framework: first, unsupervised extraction of semantically coherent native feature bases from intermediate layers using an ℓ₁-sparse autoencoder; second, causal verification and feature-level saliency visualization via Grad-FAM, a novel gradient-based attribution method; third, precise and persistent feature editing directly in weight space, accompanied by derivation of a critical suppression threshold α_crit for fine-grained robustness diagnosis. This closed-loop paradigm enables end-to-end, verifiable, and controllable intervention—from feature discovery to causally grounded editing—for the first time. We validate cross-architectural generalizability, causal fidelity, and robustness of the edits on ResNet-18 and ViT-B/16, demonstrating substantial improvements in DNN transparency and controllability.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Adversarial Learning & RobustnessReasoning under Uncertainty: Causality

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Algorithmic accountability and transparency on the webSecurity and Privacy: Data transparency and provenance
📝 Abstract
Deep neural networks achieve impressive performance but remain difficult to interpret and control. We present SALVE (Sparse Autoencoder-Latent Vector Editing), a unified "discover, validate, and control" framework that bridges mechanistic interpretability and model editing. Using an $ell_1$-regularized autoencoder, we learn a sparse, model-native feature basis without supervision. We validate these features with Grad-FAM, a feature-level saliency mapping method that visually grounds latent features in input data. Leveraging the autoencoder's structure, we perform precise and permanent weight-space interventions, enabling continuous modulation of both class-defining and cross-class features. We further derive a critical suppression threshold, $α_{crit}$, quantifying each class's reliance on its dominant feature, supporting fine-grained robustness diagnostics. Our approach is validated on both convolutional (ResNet-18) and transformer-based (ViT-B/16) models, demonstrating consistent, interpretable control over their behavior. This work contributes a principled methodology for turning feature discovery into actionable model edits, advancing the development of transparent and controllable AI systems.
Problem

Research questions and friction points this paper is trying to address.

Develops a framework for interpretable control of neural networks
Enables precise weight-space interventions for feature modulation
Quantifies feature reliance for robustness diagnostics in AI systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sparse autoencoder learns unsupervised feature basis
Grad-FAM visually validates latent features in inputs
Weight-space interventions enable continuous feature modulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.