From Preimage Search To Source-Grounded Feature Inversion

📅 2026-07-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the inherent ambiguity in conventional feature inversion methods, which often fail to establish a unique correspondence between inverted outputs and original inputs due to a lack of sample specificity. To overcome this limitation, the authors propose a source-anchored feature inversion approach that explicitly links features to the local network geometry of their source inputs, eliminating reliance on generic image priors. By employing a closed-form Wiener filtering mapping to reconstruct the accompanying signal and integrating a Jacobian-vector product (JVP)-based forward consistency residual, the method achieves precise inversion within a single backward pass. The framework enables zero-intercept mappings across architectures, depths, and channels, successfully generating images that simultaneously align with both the original input and target features in both CNNs and Transformers. Validation via predicted conditional feature maps demonstrates its efficacy in revealing the true influence of internal representations on model decisions.
📝 Abstract
Interpreting a neural network requires understanding what its internal features extract from a particular input. Feature inversion seeks to express a selected feature in the input domain, but canonical iterative methods search for an input whose re-encoded representation matches the target. Because many inputs can satisfy this constraint, target matching alone does not specify the inverse associated with the sample that generated the feature. We formulate source-grounded feature inversion by conditioning the inverse on the source-local network geometry at the target-generating input. At each boundary of the computational DAG, backpropagation provides the correct reverse dependencies but transports an adjoint signal rather than an upstream-state estimate. We locally repair this signal with a closed-form matrix Wiener map from a mean-seed VJP to the upstream state, followed by a second Wiener map for the JVP forward-consistency residual, and compose the repaired states through the same DAG in one finite reverse pass. One calibrated zero-intercept map family supports new inputs, depths, channels, and channel groups across diverse CNN and Transformer architectures, tensor components, and visual distributions without query-specific optimisation. Matched target and source controls verify that each inverse depends on the selected feature and the local operators of the sample being explained, rather than a target-independent image template. Prediction-conditioned feature atlases align these visualisations with independent interventions on the corresponding internal features. Together, source-grounded feature inversion opens the model's hidden feature hierarchy to inspection at the level of individual layers and channels, linking what the network extracts from an input to the internal evidence that shapes its decision.
Problem

Research questions and friction points this paper is trying to address.

feature inversion
neural network interpretation
source-grounded inversion
preimage search
internal feature visualization
Innovation

Methods, ideas, or system contributions that make the work stand out.

source-grounded feature inversion
Wiener map
feature interpretation
neural network explainability
computational DAG
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Kaixiang Shu
Independent Researcher