Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing evaluations for medical vision-language models (VLMs), which often examine workflows, prompting strategies, and benchmarks in isolation, leading to inflated performance estimates due to conservative predictions. Focusing on chest X-ray interpretation, this work conducts reliability stress tests on the CheXagent and MedGemma model families, proposing an evaluation protocol that jointly considers prompt sensitivity and failure modes. Furthermore, it introduces a novel decision-time routing framework that dynamically escalates to multi-agent reasoning on demand. The findings reveal how model scale influences diagnostic reliability and highlight the constraints of fixed workflows. Crucially, the proposed routing mechanism significantly optimizes the cost-quality trade-off, thereby enhancing evaluation robustness prior to clinical deployment.
📝 Abstract
Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs (CheXagent, MedGemma-4B, and MedGemma-27B) are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal" predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reliability Stress Test
Decision-Time Routing
Vision-Language Models
Multi-Agent Reasoning
Chest X-ray