Beyond Benchmark Scores: Auditing Medical Vision-Language Models for Chest X-Ray Tuberculosis Screening

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过改变队列、提示、阴性谱等条件,审计了三种医学视觉-语言模型在胸透结核病筛查中的表现,揭示了基准评分的局限性。
📝 Abstract
A medical model's benchmark score does not establish that the same conclusion holds under a different evaluation. This study tests whether claims about model ranking, score reliability and screening performance survive changes in cohort, prompt, negative spectrum, specified prevalence and operating threshold. We audit three medical vision-language models (BioMedCLIP, CheXficient, and MedSigLIP) and a general-domain OpenCLIP comparator on 12,200 chest radiograph records from four datasets (Montgomery, Shenzhen, TBX11K, and VinDr-CXR). Five fixed prompt families yield 244,000 model--image--prompt scores. No model leads every cohort and reliability criterion. Prompt-family changes alter AUROC in 21 of 48 multiplicity-controlled comparisons. Replacing healthy controls with sick non-tuberculosis controls reduces AUROC by 0.075--0.306 across all four models. On VinDr-CXR, the three medical models distinguish tuberculosis from no-finding controls substantially better than from pneumonia or lung tumor; their AUROC point estimates for both named diseases fall below 0.5. CheXficient has documented VinDr-CXR pretraining exposure, which limits the interpretation of its results. Thresholds chosen for 95\% sensitivity on TBX11K training retain that constraint by point estimate in only four of sixteen target evaluations. A five-seed supervised source model reaches 0.999 AUROC on TBX11K validation but 0.629 on each of two external cohorts. Conservative exclusion of perceptual-overlap candidates narrows this gap without closing it. These retrospective, single-task results show that discrimination, score reliability and threshold retention support different portability claims. Evidence for chest X-ray tuberculosis screening should identify the complete evaluation specification rather than attribute clinical portability to a checkpoint alone.
Problem

Research questions and friction points this paper is trying to address.

chest X-ray tuberculosis screening
model evaluation
score reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

medical vision-language models
chest X-ray tuberculosis screening
model evaluation under varying conditions
score reliability
clinical portability
M
Mushir Akhtar
Department of Mathematics, Indian Institute of Technology Indore, India
M
M. Tanveer
Department of Mathematics, Indian Institute of Technology Indore, India
Mohd. Arshad
Mohd. Arshad
Associate Professor, Department of Mathematics, Indian Institute of Technology Indore
CopulaRanking and SelectionEstimation TheoryMachine LearningData Science