Evaluating the Diagnostic Robustness of Vision-Language Models Under Visual and Textual Perturbations

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Standard accuracy metrics inadequately capture the reliability of vision-language models in medical diagnosis. This work introduces an evidence-preserving perturbation framework—encompassing slice reordering, label position swapping, and lesion removal—applied to a histopathology-validated brain MRI dataset to systematically evaluate the consistency and stability of four model families. The study reveals, for the first time, that these models are highly sensitive to both visual sequence order and textual label arrangement: reversing image sequences flips 48.9% of predictions, shuffling label positions induces 67.8% diagnostic inconsistency, and even after complete lesion removal, 76.1% of outputs remain confidently erroneous. These findings underscore hidden clinical risks masked by high nominal accuracy and advocate for a stability-centered evaluation paradigm to complement conventional accuracy-based assessment.
📝 Abstract
Standard accuracy metrics for VLMs often mask significant reliability failures in sensitive domains. In this work, we utilize a histopathology-validated brain MRI dataset to systematically assess the diagnostic robustness of four VLM families under evidence-preserving perturbations. By reordering anatomical slices and swapping target label positions, we evaluate whether models maintain consistent predictions when clinical evidence remains invariant. Our results reveal significant vulnerabilities in presentation-order stability, with models exhibiting prediction flips in up to 48.9% of cases under simple sequence reversals. We further identify a textual selection bias, where label reordering triggers inconsistent diagnoses in up to 67.8% of cases despite identical visual inputs. Negative-control tests further reveal diagnostic overcommitment: models generate categorical diagnoses in up to 76.1% of cases after expert-annotated lesion slices are removed. These results demonstrate that high accuracy can overestimate clinical reliability, masking sensitivity to sequential presentation and textual framing that is not captured by aggregate accuracy. Our findings highlight the necessity of stability-based metrics for the deployment of VLMs in safety-critical clinical applications. Our evaluation data and code will be made public upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

vision-language models
diagnostic robustness
perturbations
clinical reliability
stability
Innovation

Methods, ideas, or system contributions that make the work stand out.

diagnostic robustness
vision-language models
evidence-preserving perturbations
presentation-order stability
textual selection bias
🔎 Similar Papers
No similar papers found.