Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

📅 2026-07-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the widespread yet unverified assumption of component consistency in medical imaging AI benchmarks, which often undermines result reproducibility. Focusing on an archived chest X-ray vision-language model benchmark, the authors conduct a systematic reproducibility audit without re-invoking models or adding new annotations. Through DICOM metadata analysis, prompt binding tracing, automated label extraction, statistical code replication, and Holm-corrected multiple hypothesis testing, they uncover critical technical and metadata issues—including missing image polarity inversion, erroneous dataset splits, and truncated radiology reports. Reconstructing the evaluation cohort substantially alters the original statistical conclusions, leading to the retraction of prior performance claims and clinical assertions. The work concludes by proposing machine-verifiable control protocols to enhance the reliability of future benchmarks.
📝 Abstract
Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common cohort of 369 complete case-finding blocks changed Cochran's Q from 154.73 to 182.29. Of 45 McNemar comparisons, 27 had unadjusted p < 0.05 and 20 remained below 0.05 after Holm adjustment. These values describe only the archived automated-label matrix; they do not recover the intended prompt comparison or establish clinical performance. We withdraw the original performance, ranking, prompt-effect, and clinical claims and specify machine-verifiable controls for cohort, DICOM rendering, prompt and model identity, call status, annotation provenance, keyed analysis, and derived artifacts.
Problem

Research questions and friction points this paper is trying to address.

reproducibility
medical imaging
vision-language model
benchmarking
forensic audit
Innovation

Methods, ideas, or system contributions that make the work stand out.

forensic reproducibility audit
vision-language model
medical imaging benchmark
DICOM rendering
automated label validation
M
Mateusz Kozłowski
Independent Researcher, Kraków, Poland