CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses two critical safety risks in medical vision-language models: textual dominance over visual evidence and unwarranted confidence under insufficient evidence. We propose a training-free targeted intervention mechanism to mitigate these vulnerabilities. Methodologically, by employing dual-criteria causal tracing, temporal probing, and Tuned Lens trajectory analysis, we precisely identify that the two failure modes are mediated by spatially distinct attention heads. Targeted modulation through the removal of specific attention heads significantly reduces erroneous following behavior and restores appropriate model abstention. This work elucidates the internal causal mechanisms driving these failures within the model and validates both the effectiveness and generalizability of the proposed intervention strategy.
📝 Abstract
As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at GitHub repository.
Problem

Research questions and friction points this paper is trying to address.

Vision Language Models
Medical VQA
Arbitration Failure
Brake Failure
Mechanistic Interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Medical Vision Language Models
Mechanistic Interpretability
Causal Circuitry
Attention Head Intervention
Failure Tracing
💼 Related Jobs
No related jobs found.