AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the modality mismatch and black-box limitations inherent in existing Vision-Language Model (VLM) interpretability methods that rely on textual or internal states. We propose a training-free visual rationale generation approach that introduces a novel outer-product space mapping based on output posteriors. By executing Yes/No correlation queries over row- and column-wise banded regions of frozen model inputs, our method computes outer products to construct continuous spatial attention maps, achieving faithful and image-dependent localization without white-box access. Experiments across four VLMs demonstrate that this approach attains an AUC of 0.85, and ablating critical regions flips 53% of correct predictions, significantly outperforming existing attention mechanisms. Furthermore, the proposed method effectively mitigates model hallucinations and facilitates error correction.
📝 Abstract
When a VLM answers a visual query, current interpretability tools rely on text rationales, which use a mismatched modality, or on internal read-outs, which originate too early to reflect the final output and require white-box access to the model. We introduce AnswerMap, a training-free, task-agnostic, black-box visual rationale constructed from the output head. The image is cut into K row and K column bands, each shown alone to the frozen model along with the query in the format of a yes/no relevance question. The outer product of the row and column ``yes''posteriors gives the query-conditioned spatial map. Crucially, by defining a fixed read-out R (e.g., expectation, maximum) on top of AnswerMap, we can derive continuous outputs like location natively. This bypasses the reliance on discrete text tokens for continuous-output tasks and guarantees an image-dependent answer by construction. However, a rationale can be confabulated, so we validate AnswerMap across four models and three query distributions with two tests: (a) agreement with the model's own generated point and (b) deletion of the map's region. The map lands where the model points (AUC 0.85 against 0.38 for attention), and deleting its region flips 53% of correct answers (against 19% for attention's). Beyond establishing faithfulness, we demonstrate the map's task-agnostic utility through three distinct read-outs: its maximum flags hallucinated objects without generation, its expectation localizes correctly when the model's own pointing fails, and its top-mass region, fed back as a crop, fixes half of the model's wrong answers. AnswerMap thus offers a new lens on VLM interpretability and, through its read-outs, a new output interface for visual tasks beyond text tokens.
Problem

Research questions and friction points this paper is trying to address.

Visual Language Models
Interpretability
Spatial Rationale
Black-box
Faithfulness
Innovation

Methods, ideas, or system contributions that make the work stand out.

AnswerMap
Visual Rationale
Black-box Interpretability
Output Posteriors
Vision-Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Mohamed Eltahir
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
F
Fardows Adam
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
D
Duaa M. Tahir
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
L
Lama Alamoudi
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
S
Sana Ammar
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
A
Atheer A. Alboloshi
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
J
Jory Albluey
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia
Tanveer Hussain
Tanveer Hussain
Lecturer at Department of Computer Science, Edge Hill University
Computer VisionVideo SummarisationSaliency DetectionFire/Smoke Detection
N
Naeemullah Khan
King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia