Minimalist Explanation Generation and Circuit Discovery

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the challenge of interpreting high-dimensional image classifiers, this paper proposes an activation-matching-driven minimal explanation generation method. It employs a lightweight binary autoencoder to learn sparse input masks, jointly optimizing multi-layer feature activation alignment, output consistency, structural sparsity, and robustness. Furthermore, a circuit-readout mechanism is introduced, constructing channel-level computational graphs via forward propagation and gradient-based analysis to map input explanations to internal model mechanisms. The method generates human-readable, minimal critical-region explanations while preserving decision fidelity. Notably, it is the first to automatically discover interpretable, channel-level computational pathways in pretrained models—without architectural modification or retraining—thereby significantly improving explanation conciseness, faithfulness, and mechanistic interpretability.

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Transparent, Interpretable, Explainable MLHumans and AI: Explainable AI (XAI) for Human Understanding

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Machine learning models, by virtue of training, learn a large repertoire of decision rules for any given input, and any one of these may suffice to justify a prediction. However, in high-dimensional input spaces, such rules are difficult to identify and interpret. In this paper, we introduce an activation-matching based approach to generate minimal and faithful explanations for the decisions of pre-trained image classifiers. We aim to identify minimal explanations that not only preserve the model's decision but are also concise and human-readable. To achieve this, we train a lightweight autoencoder to produce binary masks that learns to highlight the decision-wise critical regions of an image while discarding irrelevant background. The training objective integrates activation alignment across multiple layers, consistency at the output label, priors that encourage sparsity, and compactness, along with a robustness constraint that enforces faithfulness. The minimal explanations so generated also lead us to mechanistically interpreting the model internals. In this regard we also introduce a circuit readout procedure wherein using the explanation's forward pass and gradients, we identify active channels and construct a channel-level graph, scoring inter-layer edges by ingress weight magnitude times source activation and feature-to-class links by classifier weight magnitude times feature activation. Together, these contributions provide a practical bridge between minimal input-level explanations and a mechanistic understanding of the internal computations driving model decisions.
Problem

Research questions and friction points this paper is trying to address.

Generating minimal explanations for pre-trained image classifiers decisions
Identifying critical image regions while discarding irrelevant background
Providing mechanistic interpretation of model internals through circuit discovery
Innovation

Methods, ideas, or system contributions that make the work stand out.

Activation-matching generates minimal image explanations
Lightweight autoencoder trains binary masks for sparsity
Circuit readout identifies active channels via gradients
🔎 Similar Papers
2024-06-24Neural Information Processing SystemsCitations: 13
IIT Bombay
P
P. Suhail
Department of Electrical Engineering, IIT Bombay
A
A. Anand
Department of Electrical Engineering, IIT Bombay
A
A. Sethi
Department of Electrical Engineering, IIT Bombay